Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.
317 karma · joined April 4, 2023
Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.
Very cool. Never thought the brain could have two completely different ancestors
Just because the bond markets 1/2/3/5-year bonds are bloated does that mean the bubble will pop. It just increases the risk.
I'm most curious about this sentence. What have you noticed about the similarities? I'm getting really good at asking for confidence levels, tests and pushing back, but I'm curious what you found
Is that all the science to it?
It gets worse than that though. Most harnesses that are made to handle codex and Claude cannot handle Gemini 3.1 correctly. Google has trained Gemini 3.1 to return different json keys than most harnesses expect resulting in awful results and failure. (Based on me perusing multiple harness GitHub issues after Gemini 3.1 came out)
I know they can do better
You can "feel" the llm being limited with Gemini, less so with Claude. Hopefully even less so with chatgpt