HNHacker News
TopNewBestAskShowJobs

guilamu

1,014 karma · joined October 2, 2013

submissionscomments
guilamu··on Right Click XLSX <-> CSV PowerShell Script
Extremely useful for me at work, thought I'd share: One 53 KB PowerShell file, 0 dependencies. Designed by Opus 5, two passes of bug fixes with GPT 5.6 Terra and Kimi k3 at max reasoning effort.
guilamu··on Ads and tracking infiltrated TVs. Now they're coming for monitors
Related : https://m.youtube.com/watch?v=Q9uefFYe6bM&pp=0gcJCXACo7VqN5t...
guilamu··on Israel creates fake think tank in likely attempt to dupe AI chatbots
For some reason this page exists only in French, but it says a lot about the credibility of the ADL: https://fr.wikipedia.org/wiki/Anti-Defamation_League#Controv...
guilamu··on DeepSeek V4 Flash 0731
Where does this "10x" comes from?
guilamu··on Beginning July 20, Claude Fable 5 will be included in all Max plans
FTFY: Starting July 20 Claude Fable 5 will ve excluded from Pro plans.
guilamu··on Phosh 0.56.0
https://phosh.mobi/faq/#so-what-phones-are-supported
guilamu··on Qualcomm Linux 2.0
Wouldn't you say that Valve is an exception to that rule?
guilamu··on MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
Any source to backup this claim, pretty please?
guilamu··on Oura says it gets government demands for user data
If you're concerned about that do not give internet to your tv and use any kind of tv box instead (shield tv, apple tv, etc).
guilamu··on Mistral's CEO: Europe has 2 years to stop becoming America's AI 'vassal state'
It's not. France: €0.149/kWh (~$0.175) US: ~$0.12–$0.14/kWh https://www.globalpetrolprices.com/France/electricity_prices...
guilamu··on OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
You're right, I've certainly been a bit presumptuous to call this'a benchmark'. It is indeed a flawed test. Yet,It's been giving me the occasion to try some open source models and for my workflow, some of them are incredibly competitive with sota closed source models.
guilamu··on Bob Odenkirk would like to remind you that life is a meaningless farce
Most people, including me, beg to disagree. Better Call Saul was a masterpiece.

https://www.metacritic.com/tv/better-call-saul/

guilamu··on OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
Yeah as I said this a benchmark for my usecase only, a single use case, which is obvisouly not representative of everybody's needs.

What strike me as very strange though is that 0 model were able to just use the search input already present in GravitYForms forms list page and all created a second input.

Also, I know it's not in the prompt, but adding a ctrl+f shortcut to a search input? Is that that crazy? I don't know.

guilamu··on OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
https://openrouter.ai/openai/gpt-5.5-pro

30/180 usd on Openrouter. Did I miss something?

guilamu··on OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
When nothing is noted it's max reasoning (xhigh in copilot chat in vscode if available).

The models not availble on copilot were tested through opencode (max reasoning) and deepseek v4 was tested through Cline (with max reasoning too).

guilamu··on OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
Yes those two models were tested on my own PC (local inference using my own CPU/GPU). So something my be bugged on my setup. gemma4-26b should be far better than gemma4-e4b.
guilamu··on OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
Yes, the prompt is slim by design. I might be wrong, but the point was to see what the model can do "on it's own".

The eval prompt is quite extensive: https://github.com/guilamu/llms-wordpress-plugin-benchmark/b...

guilamu··on OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
Haha, just fixed the date!

I haven't evaluated the judge benchmark. You have everything needed in the repo to do so though, so be my guest. It took me a bit of time to put all this together and won't have much more time to dedicate to it before a couple of weeks.

BTW, if you explore the repo, sorry for all the French files...

guilamu··on OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
Yes Opus 4.7 fast (no reasoning) did a worst job than Sonnet 4.6 high (with reasoning) according to Gemini 3.1 Pro evaluation.
guilamu··on OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
Just tested it on my homemade Wordpress+GravityForms benchmark and it's one of the worst model of the leaderboard performance wise and the worst value wise: https://github.com/guilamu/llms-wordpress-plugin-benchmark

I know it's only on a single benchmark, but I dont understand how it can be so bad...

guilamu··on Show HN: I blind-tested 14 LLMs on a WP plugin task. Surprising Findings
Good point! I think I won't change anything right now or I'll have to remake all tests... I'll use your input for the Level 2 task I plan on working on.
guilamu··on No more Opus for Copilot Pro plan users
Indeed. I'm still sad thought :(
guilamu··on Proton Meet isn't what they told you it was
"Proton Mail, one of the services he moved to, is ultimately controlled by the US Gov,"

Would you mind elaborating, pretty please?

guilamu··on People inside Microsoft are fighting to drop mandatory Microsoft Account
Why not just get the iso, install, activate with massgravel and be done for life?
guilamu··on People inside Microsoft are fighting to drop mandatory Microsoft Account
That's true indeed, but Microsoft is not giving us any other option so why not use the good version at home? I mean what is the risk really?
guilamu··on People inside Microsoft are fighting to drop mandatory Microsoft Account
Use LTSC. It'll fix all the issues you are mentioning here.
guilamu··on Swiftkey will soon require a Microsoft account – data to be moved to OneDrive
For those interested, I just made a quick guide to migrate from Swiftkey to Heliboard: https://github.com/guilamu/SwiftKey2HeliBoard
guilamu··on Nvidia's 10-year effort to make the Shield TV the most updated Android device
Try https://github.com/spocky/miproja1, it's awesome and will never get any ads.
guilamu··on Nvidia's 10-year effort to make the Shield TV the most updated Android device
This is Google. Just change the default launcher and you're good.
guilamu··on Latest ChatGPT model uses Elon Musk's Grokipedia as source, tests reveal
I asked 6 llms "What do you think of Grokipedia as a factual source of information?". Results: https://pastebin.com/cuxfHAr4

I then asked Claude Opus to sumup: https://markdownpastebin.com/?id=aa29d92662ac4a9ea7f9b3c1d9a...

Bottom Line All LLMs agree: Grokipedia is useful for quick orientation but unreliable for serious research, especially on political, controversial, or current event topics. Wikipedia remains the more trustworthy alternative.

Page 1 of 8Next →