964 karma · joined September 19, 2017
It’s hard to do it without killing performance and requires engineering in the DC to have fast access to SSDs etc.
Disclosure: work on ai@msft. Opinions my own.
> Our systems will smartly ignore any reasoning items that aren’t relevant to your functions, and only retain those in context that are relevant. You can pass reasoning items from previous responses either using the previous_response_id parameter, or by manually passing in all the output items from a past response into the input of a new one.
https://developers.openai.com/api/docs/guides/reasoning
Disclosure - work on AI@msft
Caching to RAM and disk is a thing but it’s hard to keep performance up with that and it’s early days of that tech being deployed anywhere.
Disclosure: work on AI at Microsoft. Above is just common industry info (see work happening in vLLM for example)
Business buyers are paying API prices, not subscription
Disclosure: Work at Microsoft on AI
https://github.com/features/copilot/cli
Disclosure: work at Msft
Explaining Why NASA's Starliner Report Is So Bad > https://www.youtube.com/watch?v=L96asfTvJ_A
OpenAI is basically ensuring that they can actually get the chips they need for the DCs they are building.
I can’t guess as to what move came first (Nvidia policy change or these DRAM deals) but I would bet this is a large if not larger factor here than “bloc my competitors.
One hack would be to use recursion and let stack exhaustion stop you.
Unless by conditionals we mean “no if/else” and not “no branch instructions”.
In other words, should the device be responsible to enforcing DRM (and more) against its owner?
System prompt: “stick to steps 1-n. Step 1 is…”
I can say confidently because our company does this. And we have F500 customers in production.
I’m just excited that our industry is lead by optimists and our culture enables our corporations to invest huge sums into taking us forward technologically.
Meta could have just done a stock buyback but instead they made a computer that can talk, see, solve problems and paint virtual things into the real world in front of your eyes!
I commend them on attempting a live demo.
> All the notable open-source frameworks implement static YaRN, which means the scaling factor remains constant regardless of input length, potentially impacting performance on shorter texts. We advise adding the rope_scaling configuration only when processing long contexts is required. It is also recommended to modify the factor as needed. For example, if the typical context length for your application is 524,288 tokens, it would be better to set factor as 2.0.
Also, if you pull too manny resources from training your next model to make inference revenue today, you’ll fall behind in the larger race.
Things they could do that would not technically contradict that:
- Quantize KV cache
- Data aware model quantization where their own evals will show "equivalent perf" but the overall model quality suffers.
Simple fact is that it takes longer to deploy physical compute but somehow they are able to serve more and more inference from a slowly growing pool of hardware. Something has to give...
Mechanistic research at the leading labs has shown that LLMs actually do math in token form up to certain scale of difficulty.
> This is a real-time, unedited research walkthrough investigating how GPT-J (a 6 billion parameter LLM) can do addition.
It’s like Mistral is choosing to fail here.
Edit: I can't even tell if its a CLI tool, an IDE plugin or a standalone IDE!
Edit 2: oh man! it's at the bottom of the page
Edit 3: "Mistral Code Enterprise is currently only available with an enterprise license." :D
Developers tend to seriously underestimate the opportunity cost of their own time.
Hint - it’s many multiples of your total compensation broken down to 40 hour work weeks.