There’s no reward for prosocial in llm rl as compared to other targets.
Humans have it since prosocial and others have evolutionary reward signals that do.
103 karma · joined September 8, 2019
There’s no reward for prosocial in llm rl as compared to other targets.
Humans have it since prosocial and others have evolutionary reward signals that do.
I've done this sort of with comfyui/same agent factory stuff, but the verification loop only works for models like fable as planner/writer, with gemini as verifier for like a very short movie. Sub 3-5 mins. After that you burn through million tokens.
Can't go too low fidelity audio/video or it craps out. Too long video and it loses consistency. Look at only snippets, it lacks global consistency, etc.
The people want better for themselves, similar to people here. Since this is speaking about people in Sulawesi, they don't really have any money or work there. It's largely jungle, mining, farms.
I'll just give an example with some of the people i spend time with whenever I'm in Indonesia.
Eko lives in his village. He largely has no real job or work. He catches fish sometimes. Different ways. Sometimes with rod. Sometimes he shimmies an air tank with an air pipe underwater and spears them. He might get bends and die of the pains of it. I told him he shouldn't fish like that, but he still does it.
You see, Eko does crazy things because wants his two kids to live better than him. He wants them to live more like those folks in Jakarta, with air conditioning and sanitary toilets. Eko thinks his life isn't great, but he makes do with it. He largely does not have a vehicle besides a weak scooter, he brings his kids together with. His hut is poor, and he built it himself with concrete and bamboo. No electrical. He defecates into a hole toilet he learned to build from his friend. Sometimes he defecates into the river.
Eko's father died recently. The father likely died of poor treatment. The hospital is about 2 hours away in the main city. The hospital is poorly run. Its red lights make you feel ill; there is barely any materials or equipment.
Sometimes he thinks he needs to bomb something to get more fish. It doesn't ever get him enough money because he needs a lot more than he does get from it. Even if his wife works on her tobacco farm, there is very little at the end of the month. Maybe his kids will die too in that same hospital. Some of my friend's kids died in similar fashion last year. Maybe an earthquake will happen again, and all the stuff that they've been doing will just be burnt away.
Poverty. It's uh, quite bleak.
Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.
I think we experimented between 4096 and 8192 to measure p99 latency performance and that was the optimal as tradeoff.
Dagger was very nice for the newer services and the ones that used straight rx were like less hell
I agree with most of this except this. Think there’s some rose tinted glasses here or I’ve got bad luck over time.
Life before ai was bad as well. There wasn’t any one to explain to you anything! You had to figure it out yourself. The people either already left or was busy with something else.
No one wrote tests (to my standard). Most of the ops works was skipped. Docs were just not there. Nobody linted properly. Just bad mannnn
Go can.
There’s no equivalent competitor, it’s the best if u want to just write lots of undifferentiated code.
To caveat this if u want to run about 50 agents or so in parallel, all the typescript projects burn ur disk via node modules. The rust ones take forever to compile and burn too much compute
There’s no equivalent competitor, it’s the best if u want to just write lots of undifferentiated code.
What’s the problem then?
Use a cheap ds4 or Luna and do the second model net net per one shot best case you save couple dollars per?
There’s a measurable performance tradeoff versus gqa so there’s reluctance.
For the most part though the new deepseek v4 tech is hca and mhc and people are still catching on like with moe and rl. Wait for 6 12 months, minimum time for next pre train.
A big problem with the existing engines like llama or sd is that they don’t support optimal graph compilation. Usually this means about a real 2 or 3x multiplier loss relative to optimal. Cuda graphs do okay but they still leave a lot on the floor
It’s usually worth it to optimize in that context if you are willing to peer into the mechanics.
Of course that’s expensive. You need to know how to appropriately pipeline and merge your kernels.
As we’ve collapsed the cost of the operations then largely the point of such an llm is constructing above.
If you’re an engineer why wouldn’t you pay more for better abstractions in an llm? That’s half the work anyways.
To put some more context into this: a manager defines the shape of a group of people. They own amorphous blob of responsibilities and various services. This leads to context confusion and lack of ownership and diffuse ability to operate services.
To fix the manager goes: Team a is responsible for say the device platform. Team b is for the applications platform. Each is responsible for their domain and the abstraction of team. Each is then responsible for their own ops, roi, quality, else. This attempts to maximize consistency of context, incentives, and scaling. Whether it works or not is frankly up in the air.
But notice this is basically the same as deciding in your intra service modularization and how you define the interfaces. Can you objectively say that whatever abstraction you usually write is correct? No. You just hope with experience and pragmatism.
The meta culture has historically been known as "move fast and break things". This seems counter to a typical cloud business.
Seeing how meta works in ads and business interfaces for whatsapp/instagram, it seems that
1. meta lacks support and customer success functions for enterprise
2. meta lacks b2b sales functions/culture
3. meta lacks a scrutable high quality operations culture for their APIs (SLAs, scaling, etc).
These aren't easy things to gain, and we see the problems that the AI providers/azure/gcp etc have had historically as consequence of these various issues.
I don’t know if I like the current world without it though.
80% of different teams code the code is poorly tested. The code doesn’t handle data consistency or asynchronous code properly because the engineers don’t know better (and frankly don’t care enough).
Dependency handling is poorly managed leading to low quality operations with improper dashboards, alarms, and ops.
Badly managed processes leads to people doing monkey work signing off checklists rather than automation.
Frankly… why is keeping any of that good? It really pisses me off seeing people accept any of that low quality but that standard is the default and not the outlier.
The project was generally a success and achieved replacement of internal bedrock services.
BUT, it came with various caveats:
1. the communication bottleneck is significant and alignment becomes even more expensive between engineers
2. engineering decisions cascade and become even more important, so you need high quality decision makers. The entire team was highly skilled, principal and distinguished engineer level.
3. every single blocker that you had administerially previously would need to be broken down, and your organization has to functionally already be in full continuous deployment mode already. if you aren't, the increased velocity is useless.
4. even with increased velocity, further bottlenecks were induced such as the rate of safe deployments through the pipeline, the ability to A/B/feature flag test, soak tests, and others.
in aws, some of the core bedrock services have been replaced with the new serving architecture. that thing was written basically with LLMs.
mind you, guy's a distinguished engineer, his team was basically all principals, but you can do it and some of the new teams are copying the style (though with less success, due to lack of technical skill).
Define the smallest market possible or something like that. I’m not sales though.
largely lags behind opus4.7/gpt5.4, but is respectable, and generally outperforms the glm/qwen equivalents anecdotally despite benchmarks.
fails to follow instructions more often, and is less code critical, but performs okay if you can decompose the task to smaller problem spaces. i.e. only do manual review, only do typechecking, only do specific component. etc
https://artificialanalysis.ai/agents/coding-agents?coding-ag...
Is it valuable to u? Is it valuable to a Chinese person? A Spaniard?
Google Translate counts as AI.
Usually it’s done in post training to enforce behavior based on prompt. Ie. System prompt with thinking:max or low or wtv.
Enforcement then goes via constrained decoding, checking for think token start and end with max lengths, or other variations
- are other teams adopting this approach? What’s the blockers if not?
- have there been problems where the models alone were not enough to debug and the devs had to fix it themselves?
- as the rate of changes has increased with more devs how have you dealt with concurrent writers with merge conflicts?
- if there was anything you could change in the approach you started with, what would it be?
Seems niche to be both uncacheable and long context?
it goes all over the place.
i'm not actually sure who your target audience is.
there's too many side tangents.
just like, structure it plz.
1. customer feels bad cuz they don't understand how llms work
2. provide high level abstracted explanation (don't dive into concepts yet)
3. provide breakdown guide of overall set of components.
4. walk through each component. don't side track. no need to explain, ROPE,GQA etc... it just distracts.
i.e. customers don't know how llms work, leading them to feel bad about their own intelligence.
at a high level llms take in words, do some math on them, and then produce words, one by one.
inside llms have these different components. we walk through them step by step.
1. tokenizer
2. embedding
3. attention
4. heads
5. ffn
6. sampling
## tokenizer
IDK i've always been able to get senior managers/managers to give me lots of leeway to do wtv i want and figure out problems. At least in most places i've worked at.
i work on platform stuff mostly, so there's always a large need for stuff. the backlog alone b4 all the agentic stuff was roughly 20-30x the capacity of any team (per year). 80% of requests were usually SVP goals and we'd just outright drop them due to lack of capacity or request HC transfer/away teams.
i.e. internal improvements alone were always massive (not gonna talk about prod code/cross team organization).
1. we need better test coverage for x,y,z
2. we need to be able to eval the long term costs (XX growth YOY, how to reduce)
3. the internal system streaming is inefficient, need to eval the alternative systems
4. we need better ops handling/management automation for issues and sev3s/sev2s. i.e. scaling, anomaly analysis, bugs introduced, improved metrics, dashboards.
5. DX stuff needs better handling, people keep confusing themselves on how to onboard. better docs on how to onboard, automation,
6. teams x,y,z are fighting with each other bcz they don't have a good grasp on systems, improve internal docs on arch and interop
7. we need automation to be able to more easily test our systems in an adhoc fashion
8. there's no linters for API platform, leading to bad results and inconsistency.
9. we're seeing bugs in the code, but aren't appropriately manualy testing after deployments. spawn 100 agents to do it, compile the results. do it every X days, feed the bugs back into the system.
i could go on and on. and this is one service, usually you own quite a few, and each one has their unique set of challenges.
For example we had this problem where we had to take in customer inputs for requests and calculate out the projected downstream TPS. This is fairly complex since we run a query parser/orchestrator.
This is expensive to write myself or to have engi's do it, but the scaling algorithms are all there and we have excel sheets for spreading out overall costs.
so then all was needed was basically write out a big spec of the reqs - give it the docs/parser code/excel sheets, then just have it span out the pieces as a sequential checklist. 1. CI/OPS 2. docs 3. test infra 4. incremental build out in phases to chain it all together.