HNHacker News
TopNewBestAskShowJobs

carterschonwald

4,018 karma · joined March 30, 2008

https://github.com/cartazio/oh-punkin-pi/blob/main/scripts/build-binary.sh <—— my open source testbed harness i mention

Personal contact first name dot lastname at gmail you know the suffix Biz email that gets auto filtered so I reply slightly sooner to them First name atsign wellposed dot sign {random gibber gabber to be ignored} com

I prefer following up email with phone because its a more efficient use of my time, :)

I'm currently building numerical computation tools for businesses using haskell.

I also like helping match awesome people with cool opportunities (when I get along with them :-) )

submissionscomments
carterschonwald··on DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
thx for the kind response!

at the very least i have tools that let me easily hit better perf for fancy dense memory layouts, and the same tooling lets me experiment with frankly wildly wacky sparse and structured memory formats. the performance claims at least on the dense side are solid so far!

the sparsity angle is because i want magic in the world. like anyone with a really chunky computer like any of those mac mini pros or serious workstation / server tier compute should be able to train from scratch their one 31b equivalent model in a week or so tops is the goal post i have in mind

carterschonwald··on DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
i definitely will be doing some drop of some faster attention kernels in the next few weeks.

like i can do all sorts of memory layout of tensors/matrices etc tricks that if you dont have the abstractions for it would just never happen. so i can optimize the kernel flops

carterschonwald··on DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
amusingly ive been working on ultra sparse llm inference/ training/ model design because nature loaths a dense graph/matrix and cause i think it shoukd be possible. i actually stood up a 20-25 percent faster than sota causal fast attention kernel yesterday, will be standing up cuda/metal/armv8 kernels too and thats gonna be fun.

i genuinely think these models should be like 0.1 percent sparse for same capabilities we associate with them today, but theres no sane way to do that with extent tools. i built the right core tech for that in 2014 when there wasnt a market, but now there is and the experimentation velocity is wild.

amusingly llms really have a hard time using my simple apis because its not in distribution array programs. but i literally stood up cpu custom memory format and micro kernel for dense causal attention in less than 24-36 hours and outperforms the equivalent fused ggml/llama cpp fast oath by like 20-25 percent

carterschonwald··on LLM Usage in Debian: Three Proposals
i think the line is: expressing that you reputationally certify its correct and its worth the time
carterschonwald··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
this mirrors my approx experience.
carterschonwald··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
i've had trouble finding any anecdotes or data about how to actually set/explore logit sampler settings
carterschonwald··on You only need the frontier model for one single edit
5.6 sol is pretty good, its definitely very very well tuned, i'm not quite sure which of the others i should use near term, but theyre doing great work
carterschonwald··on You only need the frontier model for one single edit
this so true. its really hard to make sure a model isnt going off the rails if i dont see full cot. the fact that oai and anthropic models hide it now has made them less reliable. which is a shame
carterschonwald··on Nativ: Run frontier open models locally on your Mac
its a wrapper around mlx, so thats gonna be the portability bottleneck
carterschonwald··on Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
not sure about that, but im actively working on designing ultra sparse models that i want to have perform competitively with stuff 100-10_000 times larger. ehich does yield similar throughput. time will tell id it works out
carterschonwald··on British Steel taken into public ownership to protect 'vital' UK supply
one of the most hilarious examples of how the current wh admins false protectionism, in my mind, is the crucible steel bankruptcy in jan/feb 2025.

its a technological tragedy because it was the only facility im aware of globally that could actually manufacture steel based carbide alternatives at commercial volumes. idk if the relevant equipment is being operated by anyone post bankruptcy. powder steel equipment is a bit less destroyed when turned off, but i think the key blocker is that heat cycling a 3k centigrade furnace will age the material and cause cracking thatmakes it hard to resume the powederization flows

carterschonwald··on Reducing Doom Loops with Final Token Preference Optimization
ive had doom loops on release day with opus 4.6. quantization aint the culprit ;)
carterschonwald··on Precursor
even before the llm era sites would flag me as a bot for opening 15 links to read later. its fucking infuriating now
carterschonwald··on Reducing Doom Loops with Final Token Preference Optimization
this is pretty cool. i think part of the root cause is current rlhf post training design around confidence and optics rather than cooperative transparent honesty. though its kinda an expensive hypothesis to dig into as a private individual
carterschonwald··on The Log is the Agent
This paper points at an idea, but its really only legible if you have a more developed version of the idea already. I really should write more
carterschonwald··on The bottleneck might be the air in the room
i literally had a co2 sensor for my engineering team last fall cause the space was so poorly ventilated. just measuring it continuously radically changed how everyone approached using the space packing wise and ventilation. smelled better too :p
carterschonwald··on US Supreme Court rules geofence warrants require constitutional protections
good. Of course the precise language of the ruling matters, but good.
carterschonwald··on Qualcomm to Acquire Modular
so the most notoriously patent oriented tech firm is buying this up. lol ;)

good for the founders. also explains why my resume got dropped on the floor as a desk reject :p

carterschonwald··on Prompt Injection as Role Confusion
.... i thought this was more widely known, granted i did write up a pretty wacky doc explaining way more fun experiments than these, and i have a fix that even prevents role collapse in my harness on github
carterschonwald··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
https://github.com/cartazio/oh-punkin-pi/blob/main/scripts/b...

use that to install assuming you have whatevers needed to set up.

ive a much fancier next gen thing i hope to make available as a sass mid to late summer, but if you have any feedback or questions on mine do reach out

carterschonwald··on Big Tech is borrowing like never before
“sit with” is way over used lately
carterschonwald··on I restarted a 10 year old Xeon 174 times to delete 12 flags and gain 4 tps
“worth sitting with ” is in the way to overused scale :(
carterschonwald··on Is AI ruining our skills? Early results are in – and they're not good
whats your fave/best one shot code gen?

mine def has to be the initial impl of my python to my personal fave compiler ir. i got claude chat to write it in two sentences. 4 turns helping it rmeemver where it put shit becsuse of transcript issues

carterschonwald··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
my heavily modified test bed for of oh my pi fixes this
carterschonwald··on DeepSeek Introduces Vision
gemini models are also fantastic at understanding non spoken sounds
carterschonwald··on GLM-5.2 is the new leading open weights model on Artificial Analysis
this helps so much. i do it too. with some of the newer frontier models its unclear if you can even turn it off in the first party chat apps. havent compared api semantics yet.
carterschonwald··on Claude Fable 5
its part of making sure the model actually engages in emotive communication, if i'm inventing insults i've never even thought about, i'm furious :)

saying i'm "furious" has lower entropy that incredibly implausible abuse. In some first party harnesses it just results in doom loops, but thats usually because the COT is hidden after the immediate turn in those setups. COT persistence helps with a lotta stuff

carterschonwald··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
thats a harness issue not a model issue. eg i have my own reasoninf harness that forced persisted cot
carterschonwald··on Don't trust large context windows
context window size isnt quite the issue though, its that the attention mass kinda spreads out too much and everything kinda converges to a sortah global average region full of what we know to be slop! theres some really cool ways at the harness or model layer to mitigate this. just isnt really prioritized by the labs often.
carterschonwald··on AI OSS tool repo goes archived over night after raising $7.3M Seed
i had the very strange experience last week of a recruiter listing your org as an example client, and when i looked stuff up i saw the current state.
← PreviousPage 2 of 34Next →