1,211 karma · joined March 3, 2015
> I'd also be interested in how post-processing fits in with this.
I think screenspace effects like film grain and tonemapping are excluded in the same way UI elements are rendered separately from the game.
> The inference interface uses the engine-rendered RGB image as a dense, registered observation of visible scene appearance. It provides dense, pixel-aligned evidence for object support, occlusion boundaries, composition, and local material properties; engine motion vectors separately provide temporal correspondence.
> Existing image generative models commonly rely on text embeddings, exemplar images, or spatial control fields such as depth, edges, segmentation, and pose [...] These conditions are effective for general-purpose generation and editing, but they do not uniquely determine the object identities, materials, visibility relationships, lighting decisions, and pixel-aligned detail contained in an engine-rendered frame. DLSS 5 is therefore conditioned on the rendered frame itself.
> Isn't the mote that Nvidia has is they work with studios to generate the training data from the game, then they ship a model per game?
That was true for the very first version of DLSS, from DLSS 2 on the models have been universal - the per-game adjustments are done on the inference end by changing the effect intensity or masking out objects
They have a technical report on the neural rendering part of DLSS 5 which goes into it: https://research.nvidia.com/labs/adlr/files/DLSS5_Report.pdf
(Generally I think it is alright to interpret "can't do" in terms of "the common workflow that current models will afford" and not "definitionally incapable of")
I think the field deserves more credit than that, there are plenty of interpretability tools like
* natural language autoencoders for explanations of activations: https://transformer-circuits.pub/2026/nla/index.html (demo at https://www.neuronpedia.org/llama3.3-70b-it/nla )
* easier-to-interpret language model families like Backpack models: https://aclanthology.org/2023.acl-long.506/
* attribution graphs to trace internal reasoning steps: https://www.anthropic.com/research/open-source-circuit-traci... (demo at https://www.neuronpedia.org/gemma-2-2b/graph)
* functional analyses which have identified how LLMs do arithmetic - https://arxiv.org/html/2502.00873v1 - and how refusal happens: https://arxiv.org/abs/2406.11717
* data attribution methods linking training data to specific attention heads https://arxiv.org/abs/2601.21996
If we could give a comprehensive and global explanation of an LLM's behavior in a single paragraph, we wouldn't need the model to begin with, but that doesn't mean there's absolutely no understanding of the model internals whatsoever
(Have also heard about people - not just scammers/spam operations - using generated pictures and descriptions for dating profiles and real estate ads, and I cannot understand what outcome they expect when someone follows up and immediately realizes what's up)
> - how many people in a team do you need and how long would it take to write a browser engine from scratch
The boring not-really-an-answer is "it depends on which sites you want to render properly": a browser capable of rendering Hacker News is a perfectly fine one-person project (which the book would get you to), supporting things like banking and online mail requires not just developers but also people willing to "dogfood" test it as their daily browser
In the case of TCRF, agents were also actively ignoring robots.txt and disregarding the site's instructions for interacting with it appropriately (with the "delete all data" instruction being specifically hidden from the discoverable page) - so I can't really see the case for comparing it to malware that deliberately tricks users into running it
AI does make cybersecurity more difficult, and that also goes for people using coding harnesses, sandbox your projects!
That said, what kind of errors are you getting? I think the main practical difference between Tesseract's LSTM-based OCR and newer VLM-based OCR like PaddleOCR [1] is (hopefully) getting to skip making heuristics for the layout of the document, but errors possibly compounding over multiple tokens - are you getting individual illegible words or having the documents smooshed together because the OCR can't parse the layout?
I'm less interested in the philosophy than the practical "how directly does training data influence model behavior" question, however PleIAs' SYNTH dataset [1] used to train the Baguettotron model has a small set of "self-awareness about training condition" documents which are amplified to (attempt to) give the expected answers to "what are you?" questions, filtering out any rows with "Pleias self-knowledge" under "query_seed_url" might be a good first stab at such a dataset
There's a RAM crisis on, you know, let me save you the compute!
I did kind of expect the loss of SIMD to make it paradoxically slower than the floating point version, I am curious what kind of equivalent integer/fixed-point tricks exist!
> Also, the original Box3D library in floats is guaranteed deterministic across ARM and x64 already
Oh, that rocks!
When and by whom?
> They're about as "far right" as the 80's Thatcher government.
Would you describe Thatcher passing Section 28 (which criminalized "promotion of homosexuality") as plain "conservatism" or is "far right" an appropriate descriptor in that case?
Yeah, even if you want to ignore the "political commentary" - people are correctly wary of Anthropic downgrading people or silently manipulating responses if they think you're doing distillation, why would you stake your business on someone who has repeatedly and famously done the same thing many times in a much dumber fashion?