I Use AI 100 Times per Hour
tomtunguz.com
tomtunguz.com
2:00am wake up
2:05am cold shower
2:15am breakfast: almonds, breast milk bought off facebook, 50mg adderal
2:30am begin transcribing thoughts into the machine> Within the last 24 months it’s clear that AI has become an essential coworker
I'm curious why the personification of these tools? Like, with the same logic we could call our dishwashers and washing machines employees and co-workers.
My guess is it is some form of hype-speak to raise the level of perceived importance of the technology for financial gain.
Polar opposite of what I want to happen to me.
There's joy in knowing one's tools.
Working at a higher level of abstraction while loosing knowledge at the lower level also means by some degree that one is going to be reliant on the abstraction without any understanding what is abstracted away.
My workflow involves setting up a whisper server, downloading the Whispering(1) app on my computer, and binding it to a shortcut on my keyboard and mouse. Whenever I want to write something down, I just hit the shortcut, dictate and it transcribes instantly. With a Nvidia GPU (1070 in my case), transcription is nearly instantaneous. Although I haven’t set it up on my MacBook yet, I suspect it will be just as fast with Apple Silicon
(1) https://github.com/braden-w/whispering/
You can also use an API like grok, but I'm generally wary of such services.
I'm a bit of an introvert, so I found talking out loud to be awkward at first. But now I can't go back to regular typing, given the efficiency gains.
However, ignoring the effort undercuts the distance these products have to go. Speech is great because there is just a single interface to integrate with. Obviously I'm biased to my employer's speech product, but I'm sure there are many.
Biggest thing for me was when I saw that the lem editor[0] posted on hacker news[1] was a small editor, which has 3 top level features: common lisp API, LSP support, and copilot support.
I've installed gptel[2] in emacs, and hacking up a few tools that really make it shine. Up next is figuring out voice + AI + emacs :)
[0]: https://lem-project.github.io/
> I can generate several hundred lines of code in 5-10 minutes. With the newer models, I expect this to collapse to 1-2 minutes.
It seems main usage is to make plots with R... Which is funny as the bar chart on the page can be done in Excel/Numbers/Sheets by entering a few numbers and headers.
Below 30% for all modern models.
Artificial Analysis says Whisper 3 has a 10% word error rate, although it typically does better than that in my experience. When I use Whisper 2 in the ChatGPT app it usually gets 2 or 3 words wrong per paragraph of prompt.
There is only so much AI can do because it currently lacks in certain domain knowledge.
The worst one I ever had to fix was ESPN captions of commentary for some indoor motorcross thing with dirt bikes going around a track. First, the motorbike noise, but secondly, the commentators were using the (well known to fans) nicknames for all the riders, which the AI had no idea how to transcribe, no idea who they were, and were almost impossible for me to even Google.
What is generally called AI in common speech is pretty much a class of non-linear statistical models which require some training to generate weights which are than used to fit the model. Most people that know anything about statistics knows this, so it is fine actually. Misnomers exists in all industries and all science, and we just deal with them.
Is this supposed to be a good thing?
Before Google, information is hard to find. After Google, websites are incentivized to bury information behind ads and bullshit. After LLMs, the information is juiced out like lemonade.
I wonder what the next step in the incentive landscape is?