124 karma · joined April 16, 2020
GitHub: https://github.com/chriddyp Personal website: https://chris-parmer.com
9x cheaper than Haiku 4.5 and 2 letter grades better. It's also now the fastest model (using the default speeds, not trying any of the other models "Fast" mode) to complete the exam.
Similar ballpark to Luna in price, cost, and accuracy. These are very cheap models: $0.38 to answer 40 in-depth data analytics questions (compared to $15 for Opus 5.5 or $20 for Astra).
Overall very good at data analysis - handling all of the straightforward data analytics questions correctly. It fell short answering some of the questions that required some deeper statistical analysis like looking into other variables. In other words, it's not as persistent as other models in its analysis, which I think we'd expect from how they're positioning the model.
Compared to OpenAI: GPT-6 Luna did a bit better and was about 30% the cost of Haiku 5.5. GPT-6.1 Sol got all answers correct, but was 10x more expensive.
It's 10x cheaper than Mistral Medium 3.5 from April and goes from 58% to 74% correct. Definitely a generational shift.
It's not on the Pareto curve yet, but it's good enough for data analytics, and at this rate I suspect it'll be excellent in another few months.
Full write up: https://plotly.com/blog/mistral-large-4-plotly-data-analytic...
I'm very glad to see folks innovating in this space.
There are ways around this - the llm can always be clever by invoking tools to read the file contents in a different way than the direct file contents - but this is all to say that the agentic harness layer _does_ allow for deterministic logic in between tool output and the LLM requests.
So I'd be curious to see if encouraging certain conversational behaviors might actually improve the reasoning and maybe even drive towards consensus.
Here's an example that uses almost entirely higher level components: https://dash-bootstrap-components.opensource.faculty.ai/exam...
We've also been working on `dash.templates`, which provide opinionated, prebuilt UIs - no layout code required: https://community.plotly.com/t/introducing-dash-labs-dash-2-...
For more complex 3D objects, Dash users can use dash-vtk. This includes things like point clouds, CFD simulations, 3D mesh, or 3D images.
Falcon is open source and works without an internet connection or a Plotly Chart Studio account. Falcon wires together our graphing library plotly.js (https://github.com/plotly/plotly.js/), the plotly.js chart editor (https://github.com/plotly/react-chart-editor), Electron, and some open source NPM packages for connecting to databases.
Just FYI - As a company (Plotly), we're spending most of our development effort these days on Dash Open Source (https://github.com/plotly/dash) and Dash Enterprise (https://plotly.com/dash). Truth be told, we found that most companies we worked with preferred to own the analytical backend. We also heard many stories of organizations running into roadblocks with off-the-shelf SQL or BI tools (Falcon included!). Our approach with Dash is to provide the visualization and application primitives so that you could build your own tailor-made dashboards, analytical apps, or yes, even SQL editors.
If you want to read more about where we're at, here's an essay we wrote last week on Dash: https://medium.com/plotly/dash-is-react-for-python-r-and-jul...