You can publish any publicly readable sheet/form (append rows with publicly available form):
- sample sheet: https://veneer.leftium.com/s.1RoVLit_cAJPZBeFYzSwHc7vADV_fYL...
- sample form: https://veneer.leftium.com/g.chwbD7sLmAoLe65Z8
1,808 karma · joined August 24, 2010
leftium.com
Leftium: The Element of Creativity!
You can publish any publicly readable sheet/form (append rows with publicly available form):
- sample sheet: https://veneer.leftium.com/s.1RoVLit_cAJPZBeFYzSwHc7vADV_fYL...
- sample form: https://veneer.leftium.com/g.chwbD7sLmAoLe65Z8
Eventually, I will add a polishing step to my own https://rift-transcription.vercel.app.
Right now, you can experience what true realtime streaming transcription feels like.
I plan to add two "levels" of polishing:
- Simple deterministic text replacements will be applied to both interim and final text.
- LLM polishing will only be applied right before delivery.
- It will be possible to undo one or both polishing steps. (Actually even more fine-grained undo: at the replacement rule level).
I think it would make it feel even faster.
> the UX difference between streaming and offline STT is night and day. Words appearing while you're still talking completely changes the feedback loop. You catch errors in real time, you can adjust what you're saying mid-sentence, and the whole thing feels more natural. Going back to "record then wait" feels broken after that.
If you take a lot of screenshots, I highly recommend https://shottr.cc (nagware/freemium)
- Shows preview with buttons to copy to clipboard and/or save to file
- Can be configured to automatically copy/save (open app to preview last capture)
- Preview has tons of useful features like crop, annotations, color picker, ruler, OCR
- Source mode is much better, but should use a uniform monospace font for everything. (For horizontal alignment of tables and code)
- Code blocks in WYSIWYG mode should also be monospace.
- There is a weird issue where scrolling the mouse wheel often results in the app stuck slowly scrolling up. I have a feeling it is related to the scroll sync. Happens when I scroll from either preview or source.
- brew install is more convenient, but I still had to manually allow "dangerous app." If you register with the official https://github.com/Homebrew/homebrew-cask, your app will be both listed on https://formulae.brew.sh and signed so macOS no longer considers it "dangerous." (You can also just sign your app, but I heard this is another method)
I'll probably keep using my other markdown editors/viewers, but I'll keep MarkNote around for when I need to search through my MD files.
An example of this is: I had Claude analyze the hourly precipitation forecasts for an entire year across various cities. Claude saved the API results to .csv files, then wrote a (Python?) script to analyze the data and only output the 60-80% expected values. So this avoided putting every hourly data point (8700+ hours in a year) into the context.
Another example: At first, Claude struggled to extract a very long AI chat session to MD. So Claude only returned summaries of the chats. Later, after I installed the context mode MCP[1], Claude was able to extract the entire AI chat session verbatim, including all tool calls.
1. Sometimes?
2. Described above. I also built a tool that lets the dev/AI filter (browser dev console)logs to only the loggs of interest: https://github.com/Leftium/gg?tab=readme-ov-file#coding-agen...
3. It would be interesting to combine your log compression with the scripting approach I described.
I recall I enjoyed Hoplite, Data Wing, Mini Metro, Super Mario Run, and a few others.
---
You probably already know Apple arcade curates a set of games. Many of the 'plus' versions of games have the ad/loot box features stripped or set to "free."
- make the game as functional as possible: as in the game state is stored in a serializable format. New game states are generated by combining the current game state with events (like player input, clock ticks, etc)
- the serialized game state is much more accessible to the AI because it is in the same language AI speaks: text. AI can also simulate the game by sending synthetic events (player inputs, clock ticks, etc)
- the functional serialized game architecture is also great for unit testing: a text-based game state + text-based synthetic events results in another text-based game state. Exactly what you want for unit tests. (Don't even need any mocks or harnesses!)
- the final step is rendering this game state. The part that AI has trouble with is saved for the very end. You probably want to verify the rendering and play-testing manually, but AI has been getting pretty decent at analyzing images (screenshots/renders).
# Here is an example of a simple game developed with functional architecture: https://github.com/Leftium/tictactoe/blob/main/src/index.ts
- Yes, it's very simple but the same concepts will apply to more complex games
- Right now, there is only rendering to the terminal, but you could imagine other renders for the browser and game engines
- Everything is moving through space-time at c: c is not a limit; it's just the speed everything moves
- Things that don't appear to be moving in the physical dimensions have most or all of c spent in the time dimension
- Things that move very fast in the physical dimensions have little or none of c spent in the time dimension
- I think this is similar to your section explaining time dilation, but doesn't require rotation: https://lisajguo.substack.com/i/190415584/time-dilation
---
Other questions:
- Does this theory explain why we seem to only be able to travel through time in one direction? Why does the angle/direction of rotation (not) matter?
- Named `gg` for grep-ibility and ease of typing.
- However Claude has been inserting most calls for me (and can now read back the client-side results without any dev interaction!)
- Here is how Claude used gg to fix a layout bug in itself (gg ships with an optional dev console): https://github.com/Leftium/gg/blob/main/references/gg-consol...
---
# I've been prototyping realtime streaming transcription UX: https://rift-transcription.vercel.app
- Really want to use dictation app in addition to typing on a daily basis, but the current UX of all apps I've tried are insufficient.
---
# https://veneer.leftium.com is a thin layer over Google forms + sheets
- If you can use Google forms, you can publish a nice-looking web site with an optional form
- Example: https://www.vivimil.com
- Example: https://veneer.leftium.com/s.1RoVLit_cAJPZBeFYzSwHc7vADV_fYL...
- DEMO (feel free to try the sign up feature): https://veneer.leftium.com/g.chwbD7sLmAoLe65Z8
And it's not local (uses a cloud-based transcription API)
Also doesn't seem like it's realtime streaming, either. To get the most connected typing experience, try showing results in under a second from within the first word spoken (not after the utterance is complete)
This HN comment captures why realtime streaming is important: https://hw.leftium.com/#/item/47149479
I've also been prototyping realtime streaming transcription with multimodal input: https://rift-transcription.vercel.app
One idea I was tossing around was streaming transcription + batch re-transcription:
- Use streaming transcription, which works most of the time (for example, I've found the Web Speech API pretty good, as well as moonshine)
- If the streaming transcription was poor, select the bad part and re-transcribe with a more accurate batch transcription model.
This one is unique in that it supports iPhone. I haven't seen mobile support very often.
Despite all these apps, there are two things holding me back from using a dictation app on a regular basis:
- streaming transcription: see words in realtime
- multimodal input: mix voice with keyboard
So I started prototyping this type of realtime multimodal dictation UX: https://rift-transcription.vercel.app
This HN comment captures why streaming is important for transcription: https://hw.leftium.com/#/item/47149479
React Native itself renders JSX as native components (not a web view that renders HTML/CSS).
People conflate React with HTML because that is the most common renderer, but React can be rendered into anything.
You're still assuming people will be interested in one of your ideas. There is far from 100% chance of that.
To increase this chance closer to 100%: ask people what they are interested in. "Extract" the #1 problem shared by at least 10 people/businesses (that would be worth paying at least $50/month to fix). Then offer a solution to this problem.
> There are three types of problems: 1. hair-on-fire problem, 2. 2nd biggest problem, 3. everything else
Interestingly, MarkNote has two features I requested for Prism:
1. Search
2. Sidebar that shows directory tree with multiple documents (like mdserve)
So I tried MarkNote, but these are dealbreakers for me:
- Prefer seeing actual MD rendering vs WYSIWYG.
- The layout shifts as WYSIWYG is toggled to MD text (even though it's just that line, it's distracting/annoying)
- Tables and code blocks look weird.
- Also strange window management UI for code blocks was confusing; accidentally deleted a few code blocks. Luckily my MD file was versioned via git!
- Not sure if I want auto-save (see above)This is another local-first editor I would prefer using (no install required): https://stackedit.io
---
I also prefer installing via brew. Otherwise macOS doesn't allow you to run the app (because it's not signed?). I think homebrew signs the app for you.
---
I don't think I would have tried MarkNote if it didn't have the free tier, given other editors are sufficient for my needs. And I'm not sure if I would ever need any of the pro features.
Same thing for hacker news: https://hn.leftium.com
Same thing for bookmarks/start page: https://multi-launch.leftium.com
This one allows (dancer) friends to create/manage a web site without any programming knowledge: https://veneer.leftium.com Samples:
- https://veneer.leftium.com/s.1RoVLit_cAJPZBeFYzSwHc7vADV_fYL...
---
- All my projects are hosted on Vercel (and/or Cloudflare), within their free/hobby tiers.
- No plans to monetize any of them.
- I find it more interesting to work on projects that are used by someone, whether that is myself or others.
- These projects are for learning. I would love to make a living developing novel UX like these projects. Perhaps a future project, or through someone I meet via these projects. (I did get a GitHub sponsorship, which was partially made possible by my work on these projects.)
Although you don't have any problems with lag, it is possible to efficiently compute frecency in O(1) complexity
> But with an additional trick, no recomputation is necessary. The trick is to store in the database something with units of date...
Full details: https://wiki.mozilla.org/User:Jesse/NewFrecency#Efficient_co...
This is an example python app wrapped in a (macOS) native shell using Electrobun: https://github.com/blackboardsh/audio-tts
Can you report how well Voxtral Realtime compares to the other currently supported streaming models? https://rift-transcription.vercel.app/local-setup
- Subjectively I've found Web Speech API feels the best (accuracy/latency), followed by moonshine medium
OpenAI Realtime WS API is on the roadmap, so I might be able to compare via RIFT in the future...
I also did a survey of other in-browser transcription solutions: https://github.com/Leftium/rift-transcription/blob/main/refe...
- Notably, there is an (unrelated?) moonshine demo based on transformers.js (using WebGPU) with WASM fallback.
I think most apps that use Parakeet tend to use this version of the model?
See if Parakeet (Nemotron) still uses 4GB+ with my implementation: https://rift-transcription.vercel.app/local-setup
I tried comparing Parakeet streaming with Moonshine streaming. Moonshine is smaller, and I felt it was subjectively faster with about the same level of accuracy.
I made moonshine the default because it has the best accuracy/latency (aside from Web Speech API, but that is not fully local)
I plan to add objective benchmarks in the future, so multiple models can be compared against the same audio data...
---
I made a custom WebSocket server for my project. It defines its own API (modeled on the Sherpa-onnx API), but you could adjust it to output the OpenAI Realtime API: https://github.com/Leftium/rift-local
(note rift-local is optimized for single connections, or rather not optimized to handle multiple WS connections)
uv tool install rift-local && rift-local serve --open
This opens RIFT[1], my web frontend for local transcription with a copy button. You can also compare against Web Speech API and other models (including cloud API's).So realtime streaming seems even faster. If the transcription started out bad, you can cancel and restart within a few words, before the utterance is over.
One of the reasons for my streaming transcription app: https://rift-transcription.vercel.app
- You see results in less than a second as you talk.
My app also supports multimodal input: interleave talking with typing. (Click the Replay" button to see a color-coded demo.)
Supports local models (with a little setup: https://rift-transcription.vercel.app/local-setup)
(I also implemented the previous version with only a high-level, basic understanding of websockets: https://rift-transcription.vercel.app/sherpa)
I think of Claude like a "force-multiplier."
I have been able to implement ideas I previously gave up on. I can test out new ideas much faster.
For example, https://github.com/Leftium/gg started out as 100% hand-crafted code. I wanted gg to be able to print out the expression in the code in addition to the value like python icecream. (It's more useful to get both the value and the variable name/expression.) I previously tried, and gave up. Claude helped me add this feature within a few hours.
And now gg has its own virtual dev console optimized for interacting with coding agents. A very useful feature that I would probaly not have attempted without Claude. It's taken the "open in editor" feture to a completely new level.
I have implemented other features that I would have never attempted or even thought about. For example without Claud's assistance https://ws.leftium.com would not have many features like the animated background that resembles the actual sky color.
60 minute forecast was on my TODO list for a long time. Claude helped me add it within an afternoon or so.
Note: depending on complexity of the feature I want to add the spec varies in the level of detail. Sometimes there is no spec outside Claude's plans in the chat session.
- It was definitely not one CC session. In fact, this spec is a spin-off of several other specs on several other branches/projects.
- I've actually experienced quite the opposite: I suggest an idea for the spec and Claude says "great idea!" Then I change my mind and go in the opposite direction: "great idea!" again. Once in a while, I have to argue with Claude to get my idea implemented (like adding dependencies to parse into a proper AST instead of regex.)
- One tip: it's very useful to explain the "why" to Claude vs the "what." In fact, if you just explain the why/problem without a specific solution, Claude's suggestions may surprise you!