There's excellent documentation on the formats and how to access all the data.
202 karma · joined March 9, 2007
Currently:
* Building agentic commerce
Previously:
* data & AI @ customer support startup
* founder, AI agents for eDiscovery startup (raised pre-seed)
* founding team, chief data science guy @ insurance startup (made it to Series A before being acquired).
* manager & turning coffee into code @ NLP machine learning startup).
* founder, CTO @ sustainability startup (raised seed money)
Before That:
* Machine Learning Ph.D Student @ UC San Diego.
* Graduate Student Researcher on a National Geographic Expedition.
* Stripe CTF Finalist
* 3 time Yahoo! Hackathon winner. Global Hackathon Finalist.
[ my public key: https://keybase.io/a5huynh; my proof: https://keybase.io/a5huynh/sigs/Df-98a5Ya5JkscNzgBwy5xxR-P6vhb8Rf6aM4zhr7y4 ]
There's excellent documentation on the formats and how to access all the data.
Very entertaining to watch and explains things so even non players can understand the sheer absurdity of some of the attempts.
Toyota shipped/sold more as mentioned in the article, but those are unaffected. Likewise with the Tesla, any shipped afterward the defect was discovered are unaffected.
That was a recall from 2022 (https://www.cnn.com/2022/10/06/business/toyota-bz4x-wheel-fi...) for 260 vehicles (their model BZ4X electric SUV).
The Cybertruck recall affects 3,878 vehicles (https://www.caranddriver.com/news/a60538687/2024-tesla-cyber...).
More for personal use and not quite as polished but a decent alternative for those looking to play around with the idea locally.
So far this is being used for:
- Sales -> guiding new recruits during more complex client calls
- HR -> Capturing respones during screening interviews
If you'd like to try this out feel free to DM me or email me at andrew at sightglass.ai, we're looking for more testers!
We've been building a semantic search engine for audio content focusing on podcasts. We're getting an early version for feedback.
You can search or ask questions about a particular podcast episode or feed and get back an answer as well as link to the relevant podcasts.
We also let you follow your favorite podcasts and receive summaries in your inbox whenever a new episode comes out
It'll require a beefy GPU but I've seen some fun examples like someone training a LoRA on Skyrim books.
The latest so far would be Vicuna, whose weights were just recently release.
> We release Vicuna weights as delta weights to comply with the LLaMA model
> license. You can add our delta to the original LLaMA weights to obtain
> the Vicuna weights.
Edit: took me a while to find it, here's a direct link to the delta weights: https://huggingface.co/lmsys/vicuna-13b-delta-v0tl;dr, use: https://huggingface.co/sentence-transformers/all-MiniLM-L6-v...
You can define some basic rules & it'll go out and crawl those particular sites. Or use one that someone else has built. It can also sync with your Chrome/Firefox bookmarks. Would love feedback from folks who get a chance to use it !
[0] https://yew.rs/
[1] https://github.com/seed-rs/seed
[2] https://github.com/a5huynh/spyglass - Create your own personal search engine by crawling & indexing files/docs/websites that you want.
You create some rules for topics you want to index and it'll go out and crawl them. Searching through it is a global hotkey away.
I took the idea of adding "site:reddit.com" to your Google searches and expanded on it with the idea of "lenses" to add context to your search query and give the crawler direction in terms of what to crawl & index. This means that all queries are run locally, it does not relay your search to any 3rd-party search engine. Think of it as your personal bookcase at home vs. the Library of Congress.
It's still in a super early state but would love for people to start using it and providing some feedback and see what sort of lenses people want to build and search through!
Some details about the stack for the interested:
* All Rust w/ some HTML/CSS for the client.
* Client is built w/ yew + tauri
* Backend uses tantivy to index the web pages, sqlite3 to hold metadata / crawl queue
Thanks in advance!Historically, adjusting for inflation there are also some really big ones: https://www.fool.com/investing/general/2012/08/22/a-history-...
If you'd like a more in-depth explanation, I found this post by the former Chief Scientist of Obama For America campaign (2012) that goes into the key differences between how the data was collected & used compared to Cambridge Analytica: https://medium.com/@rayid/why-what-cambridge-analytica-did-w...
Hopefully these reading materials are still useful!
- "Science of Scientific Writing" https://cseweb.ucsd.edu/~swanson/papers/science-of-writing.p...
- "Style: Lessons in Clarity and Grace" https://www.amazon.com/Style-Lessons-Clarity-Grace-10th/dp/0...