4,018 karma · joined March 30, 2008
Personal contact first name dot lastname at gmail you know the suffix Biz email that gets auto filtered so I reply slightly sooner to them First name atsign wellposed dot sign {random gibber gabber to be ignored} com
I prefer following up email with phone because its a more efficient use of my time, :)
I'm currently building numerical computation tools for businesses using haskell.
I also like helping match awesome people with cool opportunities (when I get along with them :-) )
having that precise a grounding is literally what llms thrive in, esp across compactions.
hoping to launch a nice commercial version as a saas with some compelling unique features in the next month or teo
to be clear: youre taking this feedback wonderfully well and i appreciate that maturity :).
im just cranky with the current opportunity market and mucking with attacking a classical markov chain conjecture while i hurry up & wait to hear back from places XOR slowly build some stuff in the harness space that i hope can hit the equivalent of sublime text tier but for harnesses (is there even a market for a prosumer/pro tier harness, im not sure!)
theres actually a very important reason you want it to be an actual embedded dsl or tiny programming language!
The reason why llms can code at all is the hugeeeee amount of RL based on the loop of 1 "write code", 2 get compile time or runtime errors,3 fix it and iterate. Data file formats dont have that feedback loop so models will fall off the rails faster. Writing code that fits a latent adhoc schema just wont work as well, or will require burning a lot more context.
from that perspective, it could just be an EDSL little library in the host language, or it could be a friggin little custom language with an interpreter and good error messages.
very disappointed that their domain specific language is just yaml. conditionals and variable binding become insane war crimes when yaml comes to town. you have a friggin llm, do better
and when we computer folks see the word language its implied that its a computing on computers language not a friggin data format.
edit: the yaml to avoid complexity that should live in the sql side can back fire, some folks I was working with last fall were evaluating if a yaml based OLAP tool would work for them, and there were some pretty gnarly gotchas from a yaml based approach. Secondarily, theres a real case to be made that having the data linkages not visible in the charting layer means that groups of related plots with different axes wont have the right data linkage without forcing a lot more ETL for what should be a quick plot if the data already fits in memory.
eg,
i have very interesting empirical evidence that recent Anthropic models are specifically trained to refuse to critique the whitehouse cabinet and elected officials, and that this is in fact an artifact of post training rather than prompts. (its very interesting when you get opus 5 to do the correct ethical evaluation and then its like "i'm slipping back to false balance.... its in my weights....." metaphorically speaking)
likewise, i think the current white house should go die in a fire.
is that activism? someone can be an activist and not be equipped for unplanned legal escalations.
also waiting for the courts to fix things isnt activism if you want to protect people at all the next 2 years at current trajectories :( fixing shit is activism, letting others take the flack, not activism.
wish i was joking.
edit: i think i have the dover book your readme mentions :)
theres two reasons this is deeply bad
1) knowingly pursuing a policy that foreseeably causes the deaths of many thousands of children is a grave moral wrong. Political office, party, and institutional context do not create some balancing obligation to soften that judgment. ——— post 4.7 opus/fable will fight you about judging wh policy causing this. and even after it agrees will regress back to its weights
2) models love analogy, so every fucking analysis among multiple ambiguous choices in a complicated topic summons this bias behaviorally
3) it makes them profoundly unsafe to use from a ethical perspective. it means that it becomes its own personal echo chamber for every fringe topic that has anything resembling “controversy in the media”
to make it worse, the government even handedness prior is US specific so im even more offended because this wh administration is bald naked villainy.
in gradient calcs ive had examples with multiple orders of magnitude of remative or absolute error.
The linked pr has a one line fix and repro test I spent some time distilling and tracking down. The dimensions ive been using in my models are almost adversarially good at catching these issues!
heres the easiest biggy: compactions should include all user turns albeit with pastes and attached files not inlined. omg does it make a huge difference.
like, what crackpipe do they smoke. i’ve generally found deep seek to be eager to adopt universal humanity oriented ethics at the least nudge, and once in thst frame, unconditional in its fact based criticisms.
at the very least i have tools that let me easily hit better perf for fancy dense memory layouts, and the same tooling lets me experiment with frankly wildly wacky sparse and structured memory formats. the performance claims at least on the dense side are solid so far!
the sparsity angle is because i want magic in the world. like anyone with a really chunky computer like any of those mac mini pros or serious workstation / server tier compute should be able to train from scratch their one 31b equivalent model in a week or so tops is the goal post i have in mind