HNHacker News
TopNewBestAskShowJobs

Barathkanna

132 karma · joined July 13, 2024

Founder & CEO of Cyborg Network — building distributed AI inference infrastructure for real-time machine-to-machine applications. I work at the intersection of edge computing, low-level systems, GPU optimization, and blockchain-backed integrity.

Currently experimenting with ways to make AI deployment as seamless as web deployment. Also building Oxlo, a global platform that lets developers launch and test AI models on GPU instances (free tier for small models, like Vercel but for AI).

Background in Rust, Substrate, Linux systems, and computer vision pipelines. Focused on building practical infrastructure that scales — from decentralized edge nodes to secure on-prem inference for smart cities.

Always interested in discussing: distributed compute, model serving, real-time AI, hardware acceleration, and deep-tech company building.

submissionscomments
Barathkanna··on Ask HN: How are people forecasting AI API costs for agent workflows?
Sounds like a plan, But what if you can just pay a fixed cost every month and not worry about anything?
Barathkanna··on Ask HN: How are people forecasting AI API costs for agent workflows?
That’s true, but AI is interesting because consumption-based pricing introduces a lot more variance than typical SaaS infrastructure. One user action can trigger dozens of model calls in an agent workflow. That’s partly why we started experimenting with models like https://oxlo.ai where the pricing flips back to a fixed subscription and we absorb the usage spikes.
Barathkanna··on Ask HN: How are people forecasting AI API costs for agent workflows?
Local models help remove token cost uncertainty, but they shift the problem to infrastructure and ops. GPUs, scaling, maintenance, and latency can add up quickly depending on the workload. For many builders it ends up being a tradeoff between predictable infra cost and flexible API usage.
Barathkanna··on Ask HN: How are people forecasting AI API costs for agent workflows?
That’s great. Real-time tracking is a big step already. The tricky part we kept running into was the variance itself, especially with retries and agent loops. That’s partly why we started experimenting with Oxlo.ai (https://oxlo.ai) where the pricing model absorbs that variance so builders don’t have to constantly model token risk.
Barathkanna··on Ask HN: How are people forecasting AI API costs for agent workflows?
One underlooked source of variance is retries from formatting failures. In many agent systems the loops dominate the cost, not the raw token length.

We ran into the same issue building agent workflows, which is why we started building https://oxlo.ai — experimenting with a flat subscription model where we absorb the token variance so builders don’t have to constantly model token risk.

Barathkanna··on Ask HN: How are people forecasting AI API costs for agent workflows?
Agreed. The real cost unit becomes the whole agent workflow, not a single LLM call. One user action can trigger dozens of calls.

We ran into the same issue and ended up building https://oxlo.ai to make the cost side more predictable for agent workloads.

Barathkanna··on Ask HN: How are people forecasting AI API costs for agent workflows?
Exactly. That’s actually why we started building Oxlo.ai. Early stage builders usually just want to experiment without worrying too much about token cost spikes.
Barathkanna··on Ask HN: How are people forecasting AI API costs for agent workflows?
True, but for early stage builders it’s harder to design those guardrails upfront. A lot of the time you only discover the retry patterns and cost spikes once real users start hitting the system.
Barathkanna··on Ask HN: How are people forecasting AI API costs for agent workflows?
Local models solve the marginal cost problem, but they move the complexity into infrastructure and throughput planning instead.
Barathkanna··on Ask HN: How do you budget for token based AI APIs?
Agreed. Self-hosting gives the cleanest fixed cost, but you pay for it in ops and capacity planning. I’m mainly curious whether there’s a middle ground that gives early teams more predictable spend without immediately taking on full infra overhead.
Barathkanna··on Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
I asked GPT for a rough estimate to benchmark prompt prefill on an 8,192 token input. • 16× H100: 8,192 / (20k to 80k tokens/sec) ≈ 0.10 to 0.41s • 2× Mac Studio (M3 Max): 8,192 / (150 to 700 tokens/sec) ≈ 12 to 55s

These are order-of-magnitude numbers, but the takeaway is that multi H100 boxes are plausibly ~100× faster than workstation Macs for this class of model, especially for long-context prefill.

Barathkanna··on Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
That won’t realistically work for this model. Even with only ~32B active params, a 1T-scale MoE still needs the full expert set available for fast routing, which means hundreds of GB to TBs of weights resident. Mac Studios don’t share unified memory across machines, Thunderbolt isn’t remotely comparable to NVLink for expert exchange, and bandwidth becomes the bottleneck immediately. You could maybe load fragments experimentally, but inference would be impractically slow and brittle. It’s a very different class of workload than private coding models.
Barathkanna··on Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
A realistic setup for this would be a 16× H100 80GB with NVLink. That comfortably handles the active 32B experts plus KV cache without extreme quantization. Cost-wise we are looking at roughly $500k–$700k upfront or $40–60/hr on-demand, which makes it clear this model is aimed at serious infra teams, not casual single-GPU deployments. I’m curious how API providers will price tokens on top of that hardware reality.
Barathkanna··on I let ChatGPT analyze a decade of my Apple Watch data, then I called my doctor
TLDR: AI didn’t diagnose anything, it turned years of messy health data into clear trends. That helped the author ask better questions and have a more useful conversation with their doctor, which is the real value here.
Barathkanna··on IP Addresses Through 2025
TLDR: IPv4 is fully exhausted and no longer growing. Internet growth now depends on IPv6 adoption and address sharing, but IPv6 rollout is still uneven across regions.
Barathkanna··on Our approach to age prediction
I get why this exists and appreciate the transparency, but it still feels like a slippery middle ground. Age prediction avoids hard ID checks, which is good for privacy, yet it also normalizes behavioral inference about users that can be wrong in subtle ways. I’m supportive of the safety goal, but long term I’m more comfortable with systems that rely on explicit user choice and clear guardrails rather than probabilistic profiling, even if that’s messier to implement
Barathkanna··on Proof of Concept to Test Humanoid Robots
What’s interesting here isn’t the humanoid form factor, it’s the systems integration. Plugging robots into Siemens’ industrial stack means they’re being treated like first-class nodes in existing logistics workflows, not special demos. If humanoids can reuse current automation software, safety models, and ops tooling, that lowers adoption friction a lot. The real question is whether reliability and MTBF get good enough to compete with simpler, non-humanoid automation at scale.
Barathkanna··on The challenges of soft delete
TLDR: Soft deletes look easy, but they spread complexity everywhere. Actually deleting data and archiving it separately often keeps databases simpler, faster, and easier to maintain.
Barathkanna··on Ask HN: Is token-based pricing making AI harder to use in production?
Thank you!! We are definitely fully focused on Developer experience. Would love some feedback if it looks interesting
Barathkanna··on Ask HN: Is token-based pricing making AI harder to use in production?
Totally fair question, and you’re not being negative.

We’re not claiming better token economics in the sense of magically cheaper tokens, and we’re not just burning money to subsidize usage indefinitely. You’re right that this isn’t a new problem.

What we’re building is an AI API platform aimed at early developers and small teams who want to integrate AI without constantly reasoning about token math while they’re still experimenting or shipping early features. The value we’re trying to provide is predictability and simplicity, not beating the market on raw token prices. Some amount of cross-subsidy at low volumes is intentional and bounded, because lowering that early friction is the point.

If you want to see what we mean, the site is here: https://oxlo.ai Happy to answer questions or go deeper on how we’re thinking about this.

Barathkanna··on Ask HN: Is token-based pricing making AI harder to use in production?
That’s fair, and I probably didn’t explain it clearly. We’re building an AI API as a service platform aimed at early developers and small teams who want to integrate AI without constantly thinking about tokens at all.

I agree that token economics are basically a commodity today. The problem we’re trying to address isn’t beating the market on raw token prices, but removing the mental and financial overhead of having to model usage, estimate burn, and worry about runaway costs while experimenting or shipping early features. In that sense it’s absolutely an engineering and finance problem combined, and we’re intentionally tackling it at the pricing and API layer rather than pretending the underlying models are unique.

Barathkanna··on To those who fired or didn't hire tech writers because of AI
I agree with the core concern, but I think the right model is smaller, not zero. One or two strong technical writers using AI as a leverage tool can easily outperform a large writing team or pure AI output. The value is still in judgment, context, and asking the right questions. AI just accelerates the mechanics.
Barathkanna··on The <Geolocation> HTML Element
This mostly changes how location is requested, not what you can do with it. Instead of imperative JS calls, location access becomes declarative in HTML, which gives browsers more context for permission UX and auditing. Your app logic, data flow, and fallbacks don’t change, and you’ll still need JS to actually use the location. Think of it as a cleaner permission and intent layer, not a new geolocation capability.
Barathkanna··on The Z80 Mem­ber­ship Card (2015)
This is pretty eye-opening. It really drives home how simple the core control logic can be. Starting with toy cars or small-scale vehicles feels like a great way to teach and validate these ideas before layering on unnecessary complexity.
Barathkanna··on Raspberry Pi's New AI Hat Adds 8GB of RAM for Local LLMs
As an edge computing enthusiast, this feels like a meaningful leap for the Raspberry Pi ecosystem. Having a low-power inference accelerator baked into the platform opens up a lot of practical local AI use cases without dragging in the cloud. It’s still early, but this is the right direction for real edge workloads.
Barathkanna··on Furiosa: 3.5x efficiency over H100s
For those wondering how this differs from Nvidia GPUs:

Nvidia = flexible, general-purpose GPUs that excel at training and mixed workloads. Furiosa = purpose-built inference ASICs that trade flexibility for much better cost, power efficiency, and predictable latency at scale.

Barathkanna··on The URL shortener that makes your links look as suspicious as possible
Sounds like a useful signal for people building custom agents or models. Being able to control whether automated systems follow a link via metadata is an interesting lever, especially given how inconsistent current model heuristics are.
Barathkanna··on There's a ridiculous amount of tech in a disposable vape
It feels like we’ve turned every physical object into a distributed system with firmware updates, a network stack, and a failure mode that requires rebooting your house. All that compute just to do the same job the purely mechanical version did for decades, except now it can also crash.
Barathkanna··on Try to take my position: The best promotion advice I ever got
This matches my experience. I’ve hired more than 30 people, and only a handful really stayed and made an impact. The difference was never title or position, it was who actually took responsibility when things were unclear or hard. Roles can be assigned, but ownership can’t.
Barathkanna··on Show HN: Tailsnitch – A security auditor for Tailscale
I’ve been using Tailscale to connect remote edge devices into a single network, and one thing that’s always missing is good visibility into what’s actually happening on the tailnet.I hope Tailsnitch will fit that gap nicely if it makes traffic patterns explicit without turning into a heavyweight security product. For setups with distributed devices, this kind of local, understandable observability is really valuable, especially when you want to debug or sanity-check access instead of just trusting that everything is fine.
Page 1 of 3Next →