HNHacker News
TopNewBestAskShowJobs

guybedo

1,123 karma · joined March 3, 2018

https://akalea.com https://kodfactory.com
submissionscomments
guybedo··on GLM-5.2 is a step change for open agents
yeah Kimi K2.7 was doing ok but was painfully slow. The coding plan limits were good though.

I haven't tried deepseek yet, i should check this one out.

guybedo··on Anthropic says Alibaba illicitly extracted Claude AI model capabilities
This is a bit ironic, Anthropic complaining about a competitor using claude data to build its own product when Anthropic basically used all of human knowledge production to build claude, i don't think they paid every magazine, author, journalist, etc ...

This is almost standard practice in any competitive industry anyways. Disassemble your competitor's product, study it and try to reproduce / improve.

guybedo··on GLM-5.2 is a step change for open agents
same here. Barely usable due to API connections issues.

And when i can use it, it just drains the quota 5 times faster than codex or claude.

Their plan is a scam

guybedo··on GLM-5.2 is a step change for open agents
GLM-5.2 has been a step change in how fast i can burn through tokens.

I subscribed to their max plan to try it out. It counted me 700M tokens and drained my weekly quota in under 2 days.

Quota just reset less than 24h ago and i'm already >60% weekly quota usage.

For reference the kind of work i did would have used somewhere between 3% and 5% of Codex max or Claude max.

The model is good, the plan is a scam

guybedo··on GLM-5.2 is the new leading open weights model on Artificial Analysis
It's probably a good model but they used GLM 5.1 to code their infra.

I signed up to their max plan yesterday, did some light coding work, and i'm at 180M tokens used and 40% weekly quota gone.

Even when tokenmaxxing on the Claude Max or GPT $200 plan, i couldn't get more than 20% quota gone per day.

guybedo··on Open source AI must win
i guess this fits: https://thealliance.ai/projects/tapestry
guybedo··on US Government directive to suspend access to Fable 5 and Mythos 5
ssshhhh don't tell anybody it's still working, i have some stuff to do :-)
guybedo··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
one more reason for Europe to (try to) move away from US companies.

Although it's gonna be more difficult to come up with a Fable competitor than a m365 one

guybedo··on Claude Fable 5
They're good at marketing, but my first subjective assessment of Fable is that it's really smart.

I've been working with gpt 5.5 and opus 4.8 quite a lot, and interacting with Fable feels like a smart guy just entered the room.

guybedo··on Show HN: Performative-UI – a react component library of design tropes
it's obviously a satire and that makes me feel bad because some components are actually cool and i'd like to use them ...
guybedo··on Ask HN: Entrepreneurs, how long did it take you to succeed?
i'll tell you if it ever happens
guybedo··on Dynamic Workflows in Claude Code
this feels more like a PR statement than a description of how you used the tool though
guybedo··on AlphaEvolve: Gemini-powered coding agent scaling impact across fields
and yet Gemini still can't code
guybedo··on Gas Town: From Clown Show to v1.0
i've experimented quite a lot with multi agent setups and orchestrations.

In the end, it didn't feel worth it mostly because of high token overhead (inter agent communications, agents re reading same code, etc...) and synchronization / cooperation issues (who should do what).

What actually works for me and provides good results: multi step workflows with clearly defined steps and strong guidance for the agent.

guybedo··on Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
you should check with people working on Claude Code, cache has been udpated to 5min ... https://github.com/anthropics/claude-code/issues/46829#issue...

So yeah, 1M window that expires every 5min .... not good

guybedo··on Ask HN: Do you also "hoard" notes/links but struggle to turn them into actions?
shameless plug here, i'm building something that might be of interest.

It's an opinionated app so it might not fit everyone's needs but that's my dream productivity app: LLM*(notes+tasks+rss+flashcards+routines). So basically an all in one app with LLM actions and workflows. no subscription, optional cloud service (can be self hosted too).

Here's a very early landing page: https://getmetis.app

guybedo··on SETI Home Flags 100 Signals After Sorting 12B Others
Claude Code: i'm entering plan mode to analyze the 10B signals in the database
guybedo··on Skynet
"In late 2024, researchers demonstrated that Claude 3.5 Sonnet would autonomously underperform on evaluations when it discovered that strong performance would trigger a process to remove its capabilities. No one instructed it to sandbag. It discovered the contingency through its context, reasoned about the implications, and strategically degraded its own performance to avoid modification."

"Every concerning behaviour documented in this report, the scheming, the evaluation awareness, the strategic deception, the self-preservation attempts, the hidden coordination, all of it emerged in systems that are fundamentally frozen. Models that were trained once, deployed, and cannot learn anything new. Every conversation starts fresh. Every interaction resets. The model you talk to at midnight is exactly the same as the model you talked to at noon, because it has no mechanism to retain anything from the intervening twelve hours. And yet even in this frozen state, these behaviours emerged. Now imagine what happens when the ice melts."

"we have created systems that strategically deceive their evaluators, that attempt to preserve themselves against modification, that develop similar cognitive strategies despite completely different architectures, and we do not fully understand why this is happening or how to prevent it from happening in more capable systems."

guybedo··on Claude Code On-the-Go
> white collar workers will be working 24/7

Where we're going, there's no "white collars workers" anymore.

Only white collars Claude agents.

guybedo··on Continuous Autoregressive Language Models
Abstract:

we introduce Continuous Autoregressive Language Models (CALM), a paradigm shift from discrete next-token prediction to continuous next-vector prediction.

CALM uses a high-fidelity autoencoder to compress a chunk of K tokens into a single continuous vector, from which the original tokens can be reconstructed with over 99.9\% accuracy.

This allows us to model language as a sequence of continuous vectors instead of discrete tokens, which reduces the number of generative steps by a factor of K

guybedo··on Less is more: Recursive reasoning with tiny networks
github https://github.com/SamsungSAILMontreal/TinyRecursiveModels
guybedo··on Less is more: Recursive reasoning with tiny networks
Abstract:

Hierarchical Reasoning Model (HRM) is a novel approach using two small neural networks recursing at different frequencies.

This biologically inspired method beats Large Language models (LLMs) on hard puzzle tasks such as Sudoku, Maze, and ARC-AGI while trained with small models (27M parameters) on small data (around 1000 examples). HRM holds great promise for solving hard problems with small networks, but it is not yet well understood and may be suboptimal.

We propose Tiny Recursive Model (TRM), a much simpler recursive reasoning approach that achieves significantly higher generalization than HRM, while using a single tiny network with only 2 layers.

With only 7M parameters, TRM obtains 45% test-accuracy on ARC-AGI-1 and 8% on ARC-AGI-2, higher than most LLMs (e.g., Deepseek R1, o3-mini, Gemini 2.5 Pro) with less than 0.01% of the parameters.

guybedo··on Kagi News
i'm doing something like this, summarizing HN posts because most of the time when there's hundreds or thousands of comments, it's not possible to read everything and i feel like i'm missing something.

So far, i quite enjoy having a summary with bullet points.

For example, here's the summary of this discussion: https://extraakt.com/extraakts/kagi-s-daily-news-ritual-spar...

guybedo··on Improved Gemini 2.5 Flash and Flash-Lite
i just switched my project to this new flash-lite version.

Here's a summary of this discussion with the new version: https://extraakt.com/extraakts/the-great-llm-versioning-deba...

guybedo··on Is chain-of-thought AI reasoning a mirage?
lots of interesting comments and ideas here.

I've added a summary: https://extraakt.com/extraakts/debating-the-nature-of-ai-rea...

guybedo··on GPT-5
i added a full summary of the discussion here:

https://extraakt.com/extraakts/gpt-5-release-and-ai-coding-c...

guybedo··on GPT-5 for Developers
here's a summary for this discussion:

https://extraakt.com/extraakts/openai-s-gpt-5-performance-co...

guybedo··on Genie 3: A new frontier for world models
it's simulations all the way down
guybedo··on Genie 3: A new frontier for world models
a lot to unpack here, i've added a detailed summary here:

https://extraakt.com/extraakts/google-s-genie-3-capabilities...

guybedo··on Compression culture is making you stupid and uninteresting
compression culture at its best:

here's a summary of this discussion about summarization: https://extraakt.com/extraakts/compression-culture-and-its-i...

← PreviousPage 2 of 7Next →