DeepSeek's multi-head latent attention and other KV cache tricks
pyspur.dev
pyspur.dev
it's funny that this was clear about 5% in just due to the classic chatgpt-style format and tone
It's the Internet - we never cared about such things here. Attribution and linking, yes. "Copyright" and "authorship of original works" - are you sure you're not a legacy publisher desperate to insert itself into the free exchange of knowledge and put up a toll gate? :).
I'm joking, but only a little. Unless you actually believe LLMs sin against the Church of Intellectual Property with every token they produce, this complaint feels out of place in context of a blog post summarizing research work done in the open. There are situations in which one could try to argue LLMs violate rights of some authors, but this isn't one of such situations.
The US copyright system is not a one-line profane insult topic, to me.. we are different, yes
Back to the core issue - apparently few people took a long enough look at the article to notice it was co-written by AI; i.e. there were human editors in the loop. Sure, the format is a bit off-putting, but that's IMO mostly because nobody can be arsed to write like that, even if their own thesis supervisor told them they should, as proper structuring makes it easier for the reader to understand a complex topic.
Anyway, point is, I personally have no issue with people using AI to improve their texts - LLMs already are better at writing than most people anyway. Just as long as the saved effort is put into ensuring the content itself is valuable and communicated well.
It's a great and concise way to write!
but this is something where it's up to you to decide what you want from your ghostwriting. my comments would not a system prompt make
It seems to be a common problem; so far, I've played with Rivet, n8n, and "LLM Party" nodes for ComfyUI; they all seem to focus on everything other than allowing to conveniently loop the flows.
One of the reasons why we started building pyspur is precisely because none of the other tools support loops!
If you need more support; shoot me an email; you can find it at the bottom of the github readme.
EDIT: Just added a screenshot to the readme.
They should have also posted the PySpur pipeline, it would be interesting to see the agentic flow they used in this article. I am doing a lot of this kind of worflows manually, copy pasting stuff, I'd like to have some tools to design AI flows. PySpur looks pretty interesting visually.
I think all node-based tools should offer this kind of export, and I humbly suggest that PySpur would benefit from having it too :).
--
[0] - Right click on canvas, Workflow Image -> Export -> png.
Exactly, I'm still surprised it works so well.
Also, which formatting do you prefer? I explicitly prompted it to write everything in bullet points because I find it more digestible
More recently I prefer chain-of-thought prose to the final answer. I can trigger it with a prompt even on non-reasoning models and it's usually easier to follow.
with that said, i love the content! will be bookmarking for future reference
Glad you overall liked it!
If you want to be upfront, you should mention at the start that it's written by AI instead of showing this fake author.
This would give people the choice on whether to read it.
Putting it at the end is just to give you plausible deniability. Clearly your intention is to present this as if it was written by this Mr. Kaddour, which is a lie.
EDIT: they removed the fake author in response to this comment
The author was simply there because of the website template we used; by default, it wants you to specify an author so we did. I removed the author now, thanks for making me aware!
The place that you just removed the fake author from would be a good position I think, you could even put the logo of the AI you used where the profile picture was.
"Infinite-length sequence processing" in StreamingLLM refers to handling much longer sequences than the model's training window (e.g., millions of tokens), by combining a sliding window for recent tokens with fixed attention sinks from the start of the sequence.
I can't speak for DeepSeek, but if I had to guess, I'd say that the infinite context window isn’t practical because storing all past tokens eventually becomes too expensive.
Thanks for the post, it was an excellent read!
The template will be hugely helpful for a non-programmer like me.
As far as I know, they are the only ones using it so far
k = Wx
seeing
k = xW
is jarring. Is there a reason for using horizontal vectors? Common for data science docs?
> First token: Look at 1 token (cost: O(1^2))
Umm, is this right? There is not 1 token existing before generating the first token, so how do you look at it? AI slop?
No way to know until you painstakingly verify every single assertion that the AI made! The author of this article certainly didn't, and the content was good enough to them.
Trust me, AGI is almost there.
Yes, 80% of the post is generated, yet, I still reviewed everything!
And if I had written everything from scratch, this would have probably taken a week rather than a few hours.