HNHacker News
TopNewBestAskShowJobs

tuned

65 karma · joined December 28, 2013

www.tuned.org.uk
submissionscomments
tuned··on [dead]
trying to generate images and sounds from eigenvectors as harmonic basis
tuned··on [dead]
Measuring structural information. Check out the paper, notebooks and Python library
tuned··on Starting from scratch: Training a 30M Topological Transformer
ok, thanks. I am taking it slow then
tuned··on Starting from scratch: Training a 30M Topological Transformer
no, from my point of view is being more domain-focused instead of going full-orthogonal.
tuned··on Starting from scratch: Training a 30M Topological Transformer
right. this is a proposal that needs to be tested. I started testing it on 30M parameters then I will move to a 100M and evaluate the generation on domain-specific assisting tasks
tuned··on Starting from scratch: Training a 30M Topological Transformer
> This is obviously not powerful enough to express non-linear relationships - like graph relationships.

the distance metrics used is based on energy-informed graphs that encode energy relations in a distribution called taumode, see my previous paper on spectral indexing for vector databases for a complete roll-out

tuned··on Starting from scratch: Training a 30M Topological Transformer
also: precomputing a sparse Laplacian for N vectors at dimension D (NxD) is infinitely cheaper (if using `arrowspace`, my previous paper) than computing distances on the same full dense vectors billions of times. There are published tests that compute a Laplacian on 300Kx384 space in 500 secs on a laptop on CPU. So it is a trade-off: potentially few minutes of pretaining or hours of dot-product on dense matrices
tuned··on Starting from scratch: Training a 30M Topological Transformer
if you have a corpus of code snippets to train the manifold (Laplacian) on (and a good embedding model), it is definitely possible to try something like this.
tuned··on Starting from scratch: Training a 30M Topological Transformer
it made sense to me as it is a very simple idea I guess: causal self-attention compute QKV distances computing on the full vectors for Q,K and V; the topological transformer can provide the same computation using Q, scalar K and V. Instead of [N², N², N²] -> [N², N, N²] is used. If generation is confirmed to be on par in terms of quality, the gains are evident.
tuned··on Starting from scratch: Training a 30M Topological Transformer
it most-likely will in terms of performance as it uses 50% less memory (for sure it will at inference time that is the most used operation on web services), because it can leverage longer T and D if the design is confirmed and the quality of generation is comparable to other models. If this very basic assumption is correct, it means a lot of savings in electricity as the same GPUs can resolve more requests.
tuned··on Starting from scratch: Training a 30M Topological Transformer
Thanks to all that have read. I would be glad to answer further scoped questions on the content of the post and the paper. I answered some comments that may clarify the ideas from the redesign.
tuned··on Starting from scratch: Training a 30M Topological Transformer
the idea is to have a lot of "narrow" models to work with RAG instead of one model for all the knowledge domains or also distil the metadata that is currently in enterprise Knowledge Graphs
tuned··on Starting from scratch: Training a 30M Topological Transformer
exactly, that is the current objective. To proove that generation for a specific domain is on-par with causal attention models
tuned··on Starting from scratch: Training a 30M Topological Transformer
comparisons will be run when the quality of generation will be on pair with other available models. It is useless to have preformance if the quality is not at lease on par.

The paper runs a bench (code and bench in the paper) to compare the performance with a causal attention GPT-2 model (nanoGPT) at inference (20% faster) and at training (equivalent for T and D larger than a threshold).

tuned··on Starting from scratch: Training a 30M Topological Transformer
This is a novel re-interpretation of the Transformer, based on my previous research made with a library called `arrowspace`.

It is somehow what is called a "Grassmann-like flow" but without the Plucker embedding, or also similar to what is done in DavisTensor but relying on spectral Laplacian instead of purely geometric distances.

The problem with a lot of stuff done before is that it focuses on dense representations. This architecture is focuses on sparse representation and provides a new approximation computation based on energy-informed graphs.

tuned··on Starting from scratch: Training a 30M Topological Transformer
thanks for linking.

Yes the paper compares the new architecture (that is also a fork of my implementation of nanoGPT) with Karpathy's nanoGPT. There are also links to the code and bench used.

tuned··on Starting from scratch: Training a 30M Topological Transformer
thanks for reading. I cannot retrain an existing model as the self-attention mechanism has been completely redesigned. The Keys and Values in self-attention are stored as scalars, so a latent space with traditional weights does not make sense if used in the context of a topological transformer. The two latent spaces would be somehow equivalent eventually but they would store totally different values.
tuned··on A Rust+Burn Implementation for Nanochat
Model Architecture (gpt.rs)

Multi-layer Transformer: N stacked decoder blocks with pre-norm residual connections Rotary Position Embeddings (RoPE): Replaces learned positional encodings with rotary embeddings for better length generalization Multi-Query Attention (MQA): Reduces KV cache size by sharing key/value heads across query heads RMSNorm: Parameter-free normalization for stability (instead of LayerNorm) QK-norm: Normalizes queries and keys before attention to prevent numerical instability ReLU² MLP: Uses ReLU(x)² activation for better gradient flow on GPUs Softcap Logits: Bounds output logits using tanh(x/15)*15 to prevent extreme values

tuned··on DeepSeek-OCR Compression Meets Energy Search
In this post, I demonstrate how DeepSeek's optical compression approach—treating rendered text as a visual medium—has been replicated in Rust using `burn.dev`, and how this compression primitive unlocks a new search paradigm in arrowspace v0.18.0: energy-informed retrieval that moves decisively beyond cosine similarity.
tuned··on Ask HN: What is your current side-project?
I studying biology to grow mycelium and run data analysis on growth rates and metabolism of fungi via microscopy. https://news.ycombinator.com/item?id=27362285 Anybody biology-savvy interested in collaborating, leave a comment.
tuned··on Ask HN: Who's Looking for a Co-Founder?
Hi, I am a software engineer currently in London. Looking for new ideas to experiment new tools.
tuned··on Ask HN: Who's Looking for a Co-Founder?
here from London as well. Software engineer that would like to experiment freely with new tech stack.
tuned··on The Postmodern Family Clan
is the concept of clan by definition pre-modern?
tuned··on Ask HN: Who wants to be hired? (April 2018)
Location: Edinburgh, UK

Remote: No or partially

Willing to relocate: Yes, to London or Cambridge

Technologies: Python, HTTP, SQL (especially PostgreSQL), No-SQL (MongoDB, Redis, ...), REST, Semantic Web & Linked Data, Unit Testing, Web APIs, GIS, Functional Programming, Anything-even-Pizza-as-a-service. Very interested in testing professionally my Rust or GoLang knowledge.

Résumé/CV: https://medium.com/@lorenzogotuned https://www.linkedin.com/in/lorenzomoriondo/ https://github.com/Mec-iS

Email: tunedconsulting add_a_snail gmail add_a_domain

Interested in: Satellite data, BioTech, FinTech, Research spin-offs

Looking for: Permanent job with benefits in a well-established start-up or mid-sized mature company

tuned··on Ask HN: What are some interesting papers in CS for a beginner?
Stephen Wolfram, A New Kind Of Science

https://www.wolframscience.com/

tuned··on Island Generator
If you like this stuff, check https://github.com/Mindwerks/worldengine
tuned··on What are your rabbit holes on the internet? (For instance, HN one we all share)
Being and Time by Heidegger, a never-ending book; the same for https://en.wikipedia.org/wiki/G%C3%B6del,_Escher,_Bach
tuned··on Ask HN: How did NASA make reliable software if they didn't invent unit tests?
Asserting about involved variables each 10 lines of code (:
tuned··on Ask HN: How do you write concepts? (Structure / Tools / Howto)
About "The Design of Everyday Things", there is also a course on Udacity based on the book: https://www.udacity.com/course/intro-to-the-design-of-everyd...
tuned··on Ask HN: Sick of your bank? Want to build a open source, non-profit, online bank?
In Italy we have http://www.bancaetica.it/ it's a 'popular' bank, it means that collects money on the local district to borrow to local/italian entrepreneurs. And it's ehical, it borrows only to projects with high levels of sustainability: green economy, innovation, alternative energies etc. Now it's quite up also the concept of 'social banking' http://www.social-banking.org/the-institute/what-is-social-b...
Page 1 of 2Next →