1,312 karma · joined May 28, 2021
It's a number of things like:
1. not searching for exact quotes/text, only words in the query;
2. including other lexical forms in the results (e.g. searching for things like "how do you open the run dialog" returns results for running (the sport/exercise) as well as run (an application));
3. including synonyms, typos, etc.
It seems like while each change makes sense and can improve results in some cases, over time the combined effect has made Google's search near unusable.
The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results.
1. the linux kernel has patches created by and security vulnerabilities identified by Claude/Anthropic;
2. same with other software like SQLite and rsync.
1. a particle/anti-particle pair is created at the event horizon;
2. the particle is on the outside edge of the event horizon, so "escapes" the black hole;
3. the anti-particle is on the inside edge of the horizon, so decreases the size of the black hole due to particle/anti-particle annihilation (with a corresponding particle on the inside of the black hole).
If you're after quality then Qwen3 TTS is a very good model esp. if you take some effort to craft a voice file. It is slow, so isn't practical for real-time voices (like assistants). It can also occasionally switch to a different voice to the one provided, so you may want to break up the text being processed.
I've not yet tried other recent/recentish models.
If you are after performance then two options from older models are:
1. flite with a HTS (Hidden Markov Model) voice like cmu_us_rms (male) or cmu_us_slt (female);
2. espeak/espeak-ng with an MBROLA (an Overlapped Add model) voice (mb-us1, mb-de5-en, etc.).
Alternatively, you could try using Qwen3 TTS or over voice changing model with the CMU Arctic (http://www.festvox.org/cmu_arctic/) voice data which includes audio for the rms and slt voices among others.
If you're feeling adventurous you could also try fine tuning one of the TTS models on that data to create a custom voice, though the data is likely to be in the training data for the voices, so using an audio sample may be sufficient depending on the TTS model.
The YouTube channel https://www.youtube.com/@doranchak/videos by David Oranchak, one of the people who solved the Z340 cypher, has some more details on this as well as how the Z340 cypher was cracked.
I started with writing a correct recursive descent parser. I then extended it to detect, report, and recover from common syntax errors as I encountered them so that the parser is robust. And adding a parser test case for each of these (e.g. one test for each branch through an EBNF construction).
Some examples are:
1. missing keywords when the keyword can be detected from the current context (e.g. missing semicolon at the end of a statement);
2. using the wrong token (e.g. `:` instead of `::` in a C++ namespace qualified name);
3. detecting and ignoring whitespace in a whitespace-sensitive qualification (e.g. in XML QNames);
4. keeping in the prolog state (where functions are defined) when there are errors so that functions after the error don't get lost;
5. lexing incomplete literals like `10e` so they can be handled as integers in the parser and emitting an error for them.
When I tried using agents for something as a test it immediately started doing things instead of being part of a conversation. I want to be actively involved in the process, not sit back and let AI agents write/generate stuff that I'd have to/end up rewriting/modifying anyway when it goes against the design I have in mind.
The other thing I don't like about agents is their ability to run any command [1]. That seems like a nightmare w.r.t. the potential for leaking secrets (signing/access keys, etc.) or doing damage (deleting files, database tables, etc.).
[1] You can set the option to review every command it runs, but you're then just hand-holding the agent.
4) database synchronization
1. having more memory on the card/chip and/or faster access to that memory;
2. integrated memory and compute units optimized for matrix and vector multiply add operations;
3. optimized load circuitry to e.g. read memory in the stride and span (next row, next column) access patterns common to matrices or ensure that no/few parts of the chip are stalled waiting on data or operations to complete.
Another aspect is quantizations. These are similar to SIMD vector operations in that you are performing an operation on a block of n-bit data values at the same time, so can have optimized circuitry.
For 2 or 3 valued quantizations you can reduce various addition and multiplication operations to logic operations, avoiding circuitry for things like the half-adder, full-adder, and carry-lookahead.
Then there's adding specific circuitry for common operations such as ReLU like is done in hardware acceleration of image, video, etc. processing. There's a trade off here as optimized hardware would perform better at the specific operations but if those are too specific then they can't be used by different/newer model architectures. (Though it does make sense to try and optimize common operations/logic where possible.)
It would be interesting to see if these designs can/will benefit training as well, as that would bring down the time/cost/energy of training large models as well as making it easier for local fine-tuning.
2. Only via memory recall and even then its not stable and I don't "see it" with what my eyes are seeing/processing but more with the memory/recollection part of the brain. As such I can't visualize a unique image of that, but can construct one internally, often recalling different memories/images of the various parts of what I'm constructing.
3. No. I can't visualize (or really "see") colours on their own. When visualizing a colour I recall things that are that colour from memory (an apple, grass from a particular place, the Windows XP background, etc.). It's not that I don't know/remember that something is a specific colour, just that that part is not recalled/imagined visually. Or when I do recall colours of objects/pictures I can't hold onto that image.
When someone says "imagine a pink elephant in the room" I don't see a pink elephant, but instead have an inner dialogue to the effect of "ok, so there's a pink elephant in the corner." I may also recall an image of an elephant, but it's not pink and it's not really in the room.
When there is a concrete thing that I have a memory of ("picture a squirrel", specific things from holidays, etc.) I can see flashes of still images from those. I don't see a generic squirrel, but a specific image/memory of a squirrel I have seen.
I have general placement/shape awareness but not so much colours or other specific details. I can infer things like colour from context, or have a thought to the effect of "that building was white" but I don't really see the building as white.
For models such as text/image classifiers the outputs of the model will be a list of tags, e.g. [cat, dog, mouse].
You then run the model through your test data which has the expected output, e.g. pictures of dogs would have an expected output of [0, 1, 0]. You then compare that against the model output (e.g. [0.3, 0.8, 0.1]) and work out how "wrong" the answer is (e.g. [0.3, -0.2, 0.1]).
With this value you apply back propagation where you effectively run the model in reverse, computing the "wrongness" delta at each layer for each neuron and weights. You can numerically compute the gradients for all of these and which direction in that gradient is the right answer.
You then nudge the weights in that direction and reevaluate the model. Over repeated evaluation steps the model approaches an optimal (or locally optimal) solution.
During the training of the base models, the evaluation/scoring of the model is the next token in the training data. I.e. you evaluate the model for each token subset from [1..n] in the data and evaluate that the model responds with the n+1^th token.
I'm not sure how instruction training, etc. is done but IIUC the evaluation is not at the individual/next token prediction but is on the entire response. For example, if you are training the model to write code you could run it through a compiler or syntax checker and reward (positive score) the model if it has no errors, or punish it (negative score) if it doesn't. I'm not sure what that looks like in terms of the back propagation process.
The way to avoid this is to emphasise the qualities you do want instead of specifying those you don't. For example instead of "do not cheat" say something like "you are a model student who is moral and trustworthy" -- i.e. emphasising traits that are not associated with cheating.
This is part of how/why LLMs don't truly understand what they are doing when they have been trained on a large corpus of data.
I wonder if a way to counter this is to have things like "not bad is good", "not good is bad", etc. for various antonyms and "X is Y" for synonyms, as well as other similar constructs.
Part of software development is structuring the code to make it readable and maintainable.
Part of software development is writing suitable tests to ensure that the code does what it is required to by the design and specification.
Part of software development is ensuring that the code meets things like accessible and security guidelines.
Part of software development is choosing suitable technologies and languages.
It's also confusing that the term AI has become synonymous with LLMs as it covers other things like text prediction/correction.
So AI can be a tool or a service depending on how you are using it and what you are using it for.
The key question is how good that understanding is. For example, a model would likely have a good understanding of various named colours and hex values (e.g. from the HTML specs, X11 specs, and various colour comparison websites) such that it could reasonably correlate that to a CSS entry. It's not clear if/how well a model would identify that given an image, though it should be easy to generate a dataset of image to colour name and/or hex code for training and evaluation.
What's more interesting is whether these frontier models are at their core transformer models, whether they use residual streams to facilitate learning, and whether they are using some other as yet unpublished architecture that gives them an edge.
[1] https://www.linkedin.com/posts/rpandey1234_its-mind-bending-...
I'm not sold on the agentic workflows that AI labs are pushing and haven't yet figured out if/when/how to integrate it into my workflow.
There are two general things I'm weary of with agentic workflows compared to the chat-based workflows:
1. the agentic workflows are much more inclined to go and do the thing rather than involve you in the loop -- e.g. if you are trying to design/plan something;
2. the agents are happy to go and run any command -- you can get them to prompt to confirm the action, but they could easily do something like wipe your home directory, install a random package, or something else.
I wonder how this affects the different proto-Celtic to modern Celtic languages hypothesis. [1]
This would suggest the following language tree: Breton -> (Cornish / Welsh) -> Irish Gaelic -> (Manx / Scottish Gaelic).
It may make sense to train a frontier model on an existing architecture if the base model is not available and the instruction trained version doesn't fit with what you want. There are techniques like ablation, but those could have other effects on the model, and there can still be lingering effects of the instruction training in the model that surface less frequently (e.g. on an input not covered by the ablation training).
Otherwise, fine tuning is definitely the way to go. However, you need to be careful not to over-tune the model such that it is only tuned to the data you are training it on.