3,905 karma · joined February 25, 2011
email: hn at ianww.com
The lack of formal tooling for prompt engineering drives me bonkers, and it compounds the problems outlined in the article around correctness and chaining.
Then there are the hot takes on Twitter from people claiming prompt engineering will soon be obsolete, or people selling blind prompts without any quality metrics. It's surprisingly hard to get LLMs to do _exactly_ what you want.
I'm building an open-source framework for systematically measuring prompt quality [0], inspired by best practices for traditional engineering systems.
Editing prompts is like playing whack-a-mole: once you clear an edge case, a new problem pops up elsewhere. I'd really like to be able to say, "this new prompt performs 20% better across all our test cases".
Because I haven't found a better way, I am building https://github.com/typpo/promptfoo, a CLI that outputs a matrix view for quickly comparing outputs across multiple prompts, variables, and models. Good luck to everyone else out there tuning prompts :)
Professor Scotese was a great partner and instrumental in putting this visualization together. He's acknowledged on the site, but for those interested here is his website: http://www.scotese.com/. I believe he has a more modern iteration of the paleomap that is not downloadable on the web, but for various reasons I did not get those textures in this visualization (they didn't wrap properly iirc).
He also has a nice writeup of the methods used here: https://drive.google.com/file/d/1-q0WIa7ofISFHyBe4UxvN8DIPs8...
I'm running these tests in bulk, so I prefer to automate with the CLI, or integrate with a test framework like Jest. I think the web UI is good for tinkering, but does not fit as well into real workflows.
I built this because I'm tuning a bunch of prompts and don't have a great way to do this systematically.
This CLI tool helps you pick the best prompt and model by allowing you to configure multiple prompts and variables. It outputs "before" and "after" so you can easily compare LLM outputs side-by-side and determine if the prompt has improved the quality of each example.
Example use cases:
- Deciding whether it's worth using GPT-4 over GPT-3.5
- Evaluating quality improvements to your prompt across a large range of examples
- Catching regressions in edge cases as you iterate on your prompt
It supports a handful of useful output formats: console, HTML table view, csv, json, yaml, so you can integrate into your workflow as needed. It also can be used as a library, not a CLI.
I'm interested in hearing your thoughts and suggestions on how to improve this tool further. Thanks!
I maintain a chart generation service, QuickChart (https://github.com/typpo/quickchart), which renders millions of charts per day. The most consistent pain point for users is that charts require some programming ability, or at least a strict JSON schema.
The idea is to make chart creation more approachable. Instead of messing around with D3 or Chart.js for a one-off, you can just embed https://quickchart.io/natural/red_bar_chart in an image tag or iframe and call it a day. GPT generates a reasonable look & feel.
After you have a template that you're happy with, you can modify the chart with precision, e.g. https://quickchart.io/natural/red_bar_chart?data1=3,5,7. The idea is that you don't have to mess around with chart configs, hosting, etc.
I welcome your thoughts & feedback.
1. Using embeddings to filter context into the prompt
2. Identifying common syntax errors or hallucinations of non-existent columns
3. Flagging queries that write instead of read
Plus lots of prompt finessing to get it to avoid mistakes.
It doesn't execute the queries, yet. For an arbitrary db, it's still helpful to have a human in the loop to sanity check the SQL (for now at least).
Demo at https://www.querymuse.com/query if anyone's interested
I built this tool because I found it useful (1) as a way to learn, and (2) to start basic data analysis quickly.
It uses OpenAI's completion and embeddings API, although I might be able to move it to a cheaper model eventually.
Also worth noting that state is managed in browser localStorage. I don't store your database structure unless you explicitly save/share.
Carving out a niche is more important to me than building in public. I'm not one of those guys on Twitter making $1M/yr and I don't need a personal brand.
Good luck to Cory - I've enjoyed reading his stuff.
If you want to try GPT-3 but don't have an OpenAI API key, I've set up a quick demo here until I hit my billing cap (normally users would supply their own API key): https://arkose.pages.dev/
I wrote a very crappy web app for this back in 2012: http://keepdream.me/. It emails me every morning to ask what I dreamed and records my reply. There must be better alternatives now.
How does this work? I adapted GPlates [1], an academic project that creates desktop software for geologists to investigate plate tectonic data.
Is the geocoding accurate? Even though plate tectonic models return precise results, you should consider the plots approximate within ~100km. In my tests I found that model results can vary significantly. I chose this model because it is widely cited and covers the greatest length of time.
How should I interpret the maps/colors? The graphics that wrap the globe are provided by Dr. Christopher Scotese, a geologist who runs the PALEOMAP project. You can learn more about the project and the creation of the rasters here [2]. You might also notice some old national borders. I just work with what I can get!
Why can't it look up my location? Your location probably didn't exist at the time, geologically speaking. Try switching to closer to present day (e.g. 66 Mya)
Where are all the dinosaurs? Despite the title of this post, the visualization isn't really meant to show an exhaustive list of dinosaurs or fossils (the list doesn't even show on mobile). If you want to dig into data on fossils near you, check out the Paleobiology Database Navigator [3].
[2] https://drive.google.com/file/d/1-q0WIa7ofISFHyBe4UxvN8DIPs8...
Source code: https://github.com/typpo/whispers
Aside from obvious points of interests like housing and medical care, check out "Used cars and trucks" and the noticeable spike caused by the pandemic.
Zenysis is building a product that helps governments make data-driven public health decisions. Our work is used in developing countries to support healthcare and emergency responses for hundreds of millions of people. In the past year, we've helped governments fight epidemic outbreaks, respond to natural disasters, and allocate hundreds of millions of dollars in healthcare spending. Now we are 100% focused on helping low and middle-income countries fight COVID.
Our core product is a data integration pipeline and analytics tool. On top of this, we help build early warning systems for outbreaks, tools to flag low-quality data, and other ways to identify and visualize the most effective health interventions across entire countries.
We're looking for other mission-focused engineers who care about seeing their impact in the world and are comfortable building complex, critical systems.
Apply here: https://www.zenysis.com/#careers
I maintain QuickChart [0], an open-source web service that renders images from Chart.js configs. It is useful to people who are embedding charts in static contexts such as emails or PDFs.
A lot of people from no-code communities have reached out for help with crafting their Chart.js configs. So I decided to make this interactive chart builder that lets you create a chart template and generate dynamic charts from it.
The idea is that you get most of the flexibility of hand-coding a Chart.js config, but adding your data to it from a spreadsheet/airtable/no-code app is as easy as appending a few values to your custom chart endpoint.
[0] https://github.com/typpo/quickchart and https://quickchart.io/
The benefit to visual ETL is that non-engineers can do a lot of basic data engineering. We tie this into our more complex code-based ETL pipelines. It was a game-changer for us and helps us get a lot more done.
The reason why this map is interesting is because California has a unique property tax scheme created by Prop 13. A resident may be paying 2x-10x as much tax as their neighbor in an identical house, depending on when they bought. In extreme cases the difference can reach 100x. In the renters market, landlords may charge market rate rent but pay much less in property tax.
I wanted to see how common these disparities are and display them in a visual way.
The code is quite messy, but it's available here: https://github.com/typpo/ca-property-tax
https://quickchart.io/chart?cht=gv&chl=<DOT here>
e.g. https://quickchart.io/chart?cht=gv&chl=digraph%20MyGraph%20%7Bbegin%20-%3E%20end%7D
The service is open source: https://github.com/typpo/quickchartI posted a Show HN a while back for charting project that is also minimal (zero) dependency, in a sense: https://github.com/typpo/quickchart
It's a web service that takes a Chart.js config and renders the chart as an image. Downsides compared to this: not accessible, not as lightweight. Upsides: works in email clients (most of which disable flexbox) and other embedding situations, more customization via Chart.js