120 karma · joined November 9, 2008
I initially built it to help me read AI papers before discovering AlphaXiv but then I just kept building it for the love of the game
Might have some rough edges but one feature I'm excited about is a more personalized notebook LM style audio generation capability for any readable material
(Another reason to build this was to kick the tires on a framework I've been building to power my own apps - https://hudsonkit.com - helps me ship iOS, macOS and web apps with shared primitives)
You need rapid good enough rapid turn response + rapid good enough TTS in order to make it economically for someone who's not blending it into a more compute heavy pricing model
Have you looked into it? Are you familiar with Pipecat https://github.com/pipecat-ai/pipecat - they put out interesting demos frequently
I wonder if one way to monetize certified "skill" will be to syndicate / license your provably well software engineered / tasteful AI engineered code back to the Labs for a fee
i think it's incredible because however saturated I feel this space has become, highly likely the vast majority of the world has no idea how good models have become, let alone local models
I've added local voice support for all apps/webapps I've built this year
Local TTS is on the cusp of a breakout too, basically a year behind ASR imo in terms of adoption, understanding, size and quality
Naturally, I also added the ability to highlight, leave notes and ... ask questions (via API)
Even with incredible models, it's a decent amount of work to get this to work nicely
but i love the idea
https://x.com/usetalkieapp/status/2024342996322283734/photo/...
pretty sure it's awesome - sorry OP about mentioning another project, we're all learning here :)
Been a power user of SuperWhisper and Wispr Flow for a long time and eventually decided to unify those flows - memos & dictations, everything is a file and local first, BYOK
native app uses Parakeet (v2 or V3) on iOS
Powered by Gemini (bring your own key)
https://hooked.arach.dev/ and https://speakeasy.arach.dev/ - it's a fun setup
That way you can stay comfortably in the free plan with awesome voice messages
I have a different voice for my laptop compared to my main computer and can also pick per project
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "afplay -v 0.40 /System/Library/Sounds/Morse.aiff"
}]}],
"Notification": [
{
"hooks": [
{
"type": "command",
"command": "afplay -v 0.35 /System/Library/Sounds/Ping.aiff"
}]}]
These are nice but it's even nicer when Claude is talking when it needs your attentionEasy to implement -> can talk to ElevenLabs or OpenAI and it's a pretty delightful experience
Claude user base believes in Sunday PM work sessions
as better engineers and better designers get more leverage with lower nuisance in the form of meetings and other people, they will be able to build better software with a level of taste and sophistication that wouldn't make sense if you had to hand type everything
would probably also make sense to add quick review actions in place - like ask a question to the gitlogue tool or the author during the playback