HNHacker News
TopNewBestAskShowJobs

selvan

4,292 karma · joined April 11, 2010

Founder of CheerArena (https://www.cheerarena.com) - TV grade Live channels on Youtube/Instagram/Facebook/Twitch

http://www.github.com/selvan

submissionscomments
selvan··on OpenAI Agents API
Libraries such as agent development kit (https://adk.dev/) provides abstraction over multiple LLM vendors, long-term memory (persistance + compaction) and allow us to manage subagents & their lifecycles. Vendor neutral memory & context management is a challenge as default long-term memory uses vertext AI (gemini) in ADK.
selvan··on OpenRouter is joining Stripe
Similar to how Stripe is a middleman for payments across (fragmented) banks, they want to be a middleman for (fragmented) AI models as tokens are the new currency.

As AI agents/harness/human are spenders of tokens, enabling them to derisk from being locked to a specific model provider & allowing to (re)route to any model at anytime for better leverage in a single API, similar to how they are doing for payments.

selvan··on Ask HN: What are you working on? (August 2026)
An AI agent to convert camera roll shots into cinematic reels with narration + music. Each scene of the generated reel is editable by human.
selvan··on Magenta RealTime 2: Open and Local Live Music Models
Feature demo videos: https://magenta.withgoogle.com/mrt2
selvan··on Magenta RealTime 2: Open and Local Live Music Models
Build and play AI musical instruments on your laptop!. It is a live, interactive model that you can control with MIDI and audio, in addition to text.
selvan··on Ask HN: What are you working on? (May 2026)
Creating an AI native solution to manage workflows of my live streaming business (https://www.cheerarena.com)

Most workflow softwares are complex to extend & customize. Building an AI native, structured workflow orchestrator from scratch for agentic era.

As a starting point, have designed and implemented an AI native data store to store semantic linked structured input & output data of workflow steps/tasks. These structured input/output act as spec and guard rails for the workflow tasks.

selvan··on Run interactive commands in Gemini CLI
Thanks. Fixed it.
selvan··on Run interactive commands in Gemini CLI
From the blog " Gemini CLI spawns a new process within a pseudo-terminal in the background, leveraging the node-pty library...So how does this virtual terminal running in the background show up on your screen? Think of it like a video stream. Our new serializer takes a snapshot of the pseudo terminal at every moment—capturing every piece of text, every color, and even the cursor's position. These snapshots are then streamed to you, allowing you to see and interact with the terminal application in real-time. It's not just a stream of text; it's a live feed."

Terminal serializer code: https://github.com/google-gemini/gemini-cli/blob/main/packag...

Uses @xterm/headless npm package.

selvan··on Apps SDK
An MCP server exposes tools that a model can call during a conversation and returns results according to the tool contracts. Those results can include extra metadata—such as inline HTML—that the Apps SDK uses to render rich UI components (widgets) alongside assistant messages.

More: https://github.com/openai/openai-apps-sdk-examples?tab=readm...

selvan··on Learn Your Way: Reimagining Textbooks with Generative AI
May be personalization for narration ?. Different narration style, based on their own interest.

edit: Their demo video shows they allow learners to set different narration style based on their interest.

selvan··on Leonardo Chiariglione – Co-founder of MPEG
May be, we are couple of years away from experiencing patent free video codecs based on deep learning.

DCVC-RT (https://github.com/microsoft/DCVC) - A deep learning based video codec claims to deliver 21% more compression than h266.

One of the compelling edge AI usecases is to create deep learning based audio/video codecs on consumer hardwares.

One of the large/enterprise AI usecases is to create a coding model that generates deep learning based audio/video codecs for consumer hardwares.

selvan··on OpenAI’s Windsurf deal is off, and Windsurf’s CEO is going to Google
Cursor - co-pilot/AI pair programming usecases.

Claude Code - Agentic/Autonomous coding usecases.

Both have their own place in programming, though there are overlaps.

selvan··on WASM Agents: AI agents running in the browser
Ship AI Agents as a web page :-)
selvan··on Ask HN: What Are You Working On? (June 2025)
CheerArena - Your Own TV Grade Live Channel on Youtube

Have created a real-time media mixing mobile app that helps to setup TV grade Live channel on Youtube/Facebook/Twitch/Instagram.

Our product scales from individual to institutions, camera in mobiles to network of cameras, indoor to outdoor sports and events.

Details: https://www.cheerarena.com/

Realtime mixing studio - https://play.google.com/store/apps/details?id=com.cheerarena...

selvan··on Tracking Copilot vs. Codex vs. Cursor vs. Devin PR Performance
Total PRs between Codex vs Cursor is 208K vs 705, this is an enormous difference in absolute PRs. Since cursor is very popular, how does their PRs is not even 1% of codex PRs?.
selvan··on Making Video Games (Without an Engine) in 2025
For simpler games, libraries such as raylib or lightweight opensource game engines such as Amulet (https://www.amulet.xyz/) / Love 2D are good fit.
selvan··on Launch HN: Karsa (YC W25) – Buy and save stablecoins internationally
Curious, what would be the motivation of the sellers to trade high inflation currency?.

It make sense for buyers as they want to move to stabe currency. But how about sellers?. What are they gonna do with the high inflation currency ?.

One motivation could be of very high margin due to high risk involved.

selvan··on Ask HN: DAO for health insurance – A human collective for betterment
Fraud detection is code.

Not replacing hospitals/doctors, but replacing insurance companies.

selvan··on Gemini 2.0: our new AI model for the agentic era
Get started documentation on Multimodal Live API : https://ai.google.dev/api/multimodal-live
selvan··on FLUX is fast and it's open source
From the PDF - "One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are "search" and "learning".

The second general point to be learned from the bitter lesson is that the actual contents of minds are tremendously, irredeemably complex; we should stop trying to find simple ways to think about the contents of minds, such as simple ways to think about space, objects, multiple agents, or symmetries. All these are part of the arbitrary, intrinsically-complex, outside world. They are not what should be built in, as their complexity is endless; instead we should build in only the meta-methods that can find and capture this arbitrary complexity. Essential to these methods is that they can find good approximations, but the search for them should be by our methods, not by us. We want AI agents that can discover like we can, not which contain what we have discovered. Building in our discoveries only makes it harder to see how the discovering process can be done."

selvan··on Ask HN: What are you working on (September 2024)?
Working on - "real-time conversations in rich video streaming". Have created rich video composition, mixing, streaming studio (http://www.thecheerlabs.com), working on to bring real-time conversations that can be mixed in real-time for streaming/recording.
selvan··on Ask HN: Video editing library for browser similar to moviepy
Have used https://github.com/redotvideo

You may wanna move this post to "Ask HN:"

selvan··on An experiment in UI density created with Svelte
VS Code Editor which is based on Electron, is really fast, even with large codebase & many open tabs. Their monaco engine (https://microsoft.github.io/monaco-editor/) uses custom, virtual code processor that is optimized for surgically updating underlying DOM. It also uses WebGL + canvas rendering to show minimap of the file.

Similar approach (custom virtual processor) is leveraged by Google docs/sheets.

Canvas rendering may be the last resort when nothing worked.

selvan··on Cloudflare acquires PartyKit to allow developers to build real-time multi-user
Chat, Audio/Video Conferencing apps are other examples.
selvan··on Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
Nice, (code like) Refactoring meets speech-to-text
selvan··on Nvidia CEO Jensen Huang announces new AI chips: ‘We need bigger GPUs’
Microsoft and AWS would have a partnership with AMD/Intel for their GPUs, if those are capable and widely used as Nvidia's.

Microsoft has partnetship with OpenAI and also with Mistral.

Present convenience may not hold true in future. Nvidia knows that well.

selvan··on Sora: Creating video from text
Ad generation usecases are getting interesting with Video generation + Controlnet + Finetuning
selvan··on Peer-to-Peer decentralized networks for economic transactions
https://nammayatri.in/open/ - Raid hailing service that uses beckn

https://ondc.org/ - P2P commerce network that uses beckn

selvan··on Ask HN: How did risks of SVB's investments gone unnoticed?
Sounds like failure of financial due diligence. A proper due diligence could have uncovered the concealments.
selvan··on Ask HN: How did risks of SVB's investments gone unnoticed?
Wouldn't the details audited by accounting firms?
Page 1 of 3Next →