Throw more AI at your problems
frontierai.substack.com
frontierai.substack.com
Hell, you can even predict where by looking at the lowest-paid people in the organization. Work that isn't valued today won't start being valued tomorrow.
Sorry, not making money— I meant attracting investment.
AI tech is generally undertested and experimental, yet marketed as solid and reliable.
Raising money is about convincing investors that you can solve someone else's problem. Or even more abstractly, about convincing investors that you can convince other investors that you can convince other investors that you can solve someone else's problem. That's enough layers of indirection that signaling games start to overtake concrete value a lot of the time, and that's what gets you your Theranoses and your FTXes and the like. You get into "the market can remain irrational longer than you can remain solvent", or its corollary, "the market can remain irrational long enough for you to exit before it wakes up".
But making money means you are solving a problem for your customer, or at least, that your customer thinks you're solving a problem with no additional layers of indirection. And if we take "making money" to mean "making a profit", it also means you're solving their problem at less cost than the amount they're willing to pay to solve it. You are actually creating net value, at least within a sphere limited to you and your customer (externalities, of course, are a whole other thing, but those are just as operative in non-profitable companies).
I think this is one of the worst things about the way business has done today. Doing business, sustainably and profitably, is an excellent way to keep yourself honest and force your theories to actually hold up in a competitive market. It's a good thing for you and for your users. But business has become so much about gathering sufficient capital to do wildly anticompetitive things and/or buy yourself preferential treatment that we're losing that regulating force of honesty.
[1] see my HN profile for more on that if you care
For some AIs, if you ask them to complete it for "Java" and/or "Regexes" first, then they give realistic answers for "AI". But others (mostly, online commercial ones) are just relentlessly positive even then.
Prompting to complete it for "Python" usually remains positive though.
So one remedy is to have it just answer in plaintext, and then use a second, more specialized model that's specifically trained to turn plaintext into json. Whether this chain of models works better than just having one model all depends on the distribution match penalties accrued along the chain in between.
Also you don't need to use a model to build a json from plaintext answers lol, just use a programming language.
Now, whether hardware is cheap enough or AI is smart enough is an entirely different question...
Reading through this, I could not tell if this was a parody or real. That robot image slopped in the middle certainly didn't help.
as a general rule, virtually any analogy that involves anthropomorphizing LLMs is at best right for the wrong reasons— a stopped clock— leads to conclusions ranging from misleading to actively harmful.
For example, I want to scrape a collection of sites. The agent would at first apply the whole HTML to the context to extract the data (expensive but it works), but then there is another agent that sees this pipeline and says "hey we can write a parser for this site so each scrape is cheaper", and iteratively replaces that segment in a way that does not disrupt the overall task.
The unscalable thing is often like “buy it cheap, buy it twice” but it’s also often like “buy it cheap, only fix it if you use it enough that it becomes unsuitable”. Makers endorse both attitudes. Knowing when which applies is the challenging bit
Instead I implemented low tech “RAG” or “data source rules”. It’s a list of general rules you can attach to a particular data source (ie database). Rules are included in the generations and work great. Examples are “Wrap tables and columns in quotes” or “Limit results to 100”. It’s simple and effective - I can execute the generate SQL again my DB for insights.
The easiest example I can come up is imagine you just dump a restaurant menu into a vector DB. The menu is from a hipster restaurant and instead of having "open hours" or "business hours" the menu says "Serving deliciousness between 10:00am and 5:00pm"
Naïve RAG queries are going to fail miserably on that menu. "When is the restaurant open?" "What are the business hours?"
Longer context lengths are actually the solution for this problem, when context is small enough and the potential of ambiguity is high enough, LLMs are the better tool.
Just a reminder that smaller fine tuned models are just as good at solving the problems they are trained to solve, as large models are.
> Oftentimes, a call to Llama-3 8B might be enough if you need to a simple classification step or to analyze a small piece of text.
Even 3B param models are powerful now days, especially if you are willing to put the time into prompt engineering. My current side project is working on simulating a small fantasy town using a tiny locally hosted model.
> When you have a pipeline of LLM calls, you can enforce much stricter limits on the outputs of each stage
Having an LLM output a number from 1 to 10, or "error" makes your schema really hard to break.
All you need to do is parse the output and it if isn't a number from 1 to 10... just assume it is garbage.
A system built up like this is much more resilient, and also honestly more pleasant to deal with.
I'm a bit confused by their thinking it's a good thing while being confused about why the subject has "disappeared from the conversation".
Could anyone here shed some light / share an opinion on it/why "long context windows" aren't discussed any more? Did everyone decide they're not useful? Or they're so obviously useful that nobody wastes time discussing them? Or...
- that try to extract the factual / knowledge content and try to update the rest of the system (e.g. if the user chats you to not send notifications after 9pm, ideally you’d like the whole system to reflect that. if they say they like the color gold, you’d like the recommendation system to know that.)
- detect the emotional valence of the user’s chat and raise an alert if they seem, say, angry
- speculative “given this new information, go over old outputs and see if any of our assumptions were wrong or apt and adjust accordingly”
- learning/feedback systems that run evals after every k runs to update and optimize the prompt
- systems where there is a large state space but any particular user has a very sparse representation (pick the top k adjectives for this user out of a list of 100, where each adjective is evaluated in its own prompt)
- llm circuits with detailed QA rules (and/or many rounds of generation and self-reflection to ensure generation/result quality)
- speculative execution in order to trade increased cost / computation for lower latency. (cf. graph of thoughts prompting)
- out of band knowledge generation, prompt generation, etc.
- alerting / backstopping / monitoring systems
- running multiple independent systems in parallel and then picking the best one or even merge the best results across all of them.
the more, smaller prompts, the easier to do eval and testing, as well as making the system more parallelizable. also, you get stronger, deeper signals that communicate a deeper domain understanding to the user s.t. they think you know what you’re doing.
but the point is every bit as much that you are embodying your own human cognition within the llm— the reason to do that in the first place is because it is virtually infinitely scalable when it makes it out of your brain and onto/into silicon. even if each marginal prompt you add only has a .1% chance of “hitting”, you can just trigger 1000 prompts and voila: one more eureka moment for your system.
sure, there’s diminishing returns in the same way that the CIA wants 100 Iraq analysts but not 10000. but unlike the CIA, you dont need to pag salaries, healthcare, managers, etc. it all scales basically linearly. and, besides, is extremely cheap as long as your prompting is even halfway decent.
I recently discovered BERTopic, a Python library that bundles a five-step pipeline of now pretty old (relatively) NLP approaches in a way that is very similar to how we were already doing it, now wrapped in a nice handy one-liner. I think it's a great exemplar of the approach that will probably emerge from the hype storm on top.
(Disclaimer: I am not an AI expert and will defer to real data/stats nerds on this.)
people have this massive misconception about what LLMs can do and how to get good results out of them where they think they just kinda ask it for stuff and voila, it appears. it could not be any further from reality. it is an interactive tool that you get into the right "frame of mind" (scare quotes because this is a descriptive analogy, not one meant to convey or impart mechanical sympathy) and then ask it... literally anything you want. but you have to DISCOVER these things-- these (nearly) magic words, phrases, encodings of the problem, etc. that get the LLM to generate results in a way resonant with the way that you do / the way that you want it to. then you get the generation part "for free". you teach it to be a perfect painter, and then let it paint (shoutout Robert Pirsig // Zen and the Art of Motorcycle Maintenance).
the whole "LLM applications to <X>" stuff is such a mirage, and so tied up in a misbegotten view of how they work and what they're good at. if you have a hyper-hyper-specialized domain, sure, you might need to... find some specialized data. find or train some bespoke model. but for damn near anything written in english (can't speak to other languages) it is as good at coding as it is at bioinformatics as it is at statistics as it is at sociology as it is at literature (to be rather flippant). just do some introspection on how you do your job, how you think through problems, how you generate your output, and then "outrospect" it "into the computer". once you do that, voila: you have a virtually infinitely scalable simulacrum of your own cognition. in no way is it "plug in the ai and it will automatically take my job, one size fits all". the reason why they works so well is precisely because they are NOT that.
think about it this way: if you knew that each day when you went to sleep all of your memories were deleted, but you and your brain otherwise worked exactly the same way as yesterday, what would you do? (the concept of the mediocre but cute movie "50 First Dates") what notes would you leave for yourself to get back into the same context as you were in when you went to sleep the night before? how would you convince yourself that the artifacts you leave for yourself in the morning are true? how would you convey the subtlety of your thoughts & feelings, idiosyncracies, point of view, style & syntax? how would you grow the system over time? the LLMs are closer to speaking the language of your own thoughts than any human and every human invention in all of history. prompting is just the way to incept ideas and thoughts, systems, etc. into it's ai mind in a way that i find to be profoundly similar to how our own perception is not "reality" per se, but rather the image of reality that we see within our own minds. an image that we build up from birth as we grow to understand reality and the world around us.