As of today, Ploy’s agent runs on GPT-5.6 Sol, the flagship tier of the model family OpenAI released this morning.
Wait a moment, did they make the switch based on half a days of playing with Sol? Are these companies ran by teenagers? As of today, Ploy’s agent runs on GPT-5.6 Sol, the flagship tier of the model family OpenAI released this morning.
Wait a moment, did they make the switch based on half a days of playing with Sol? Are these companies ran by teenagers?We have been testing GPT 5.6 for about a week as a preview model through a YC relationship, providing them feedback on the model. Our evals run in github CI and we can run them all in about 15 minutes against our eval bench of 115+ web design and marketing related jobs that ploy.ai specializes in.
then after we toggled it on (through a posthog feature flag) we actively monitored for failures.
I came from running Webflow, which powers > 1% of the internet so trying my best to relay all of that knowledge to ploy to power more % of the internet!
The funny thing is - when I first saw ploy, I didn't take it very seriously since so many of the signals that used to signify quality (decent design, copy, hard technical problems) are easy to fake. Plus the "grow while you sleep" space is crowded with weak players.
I wonder what the new markers of quality will be, which would separate the hand-crafted (to the extent possible) work v/s slop.
Funnily enough, we spent a long time on our brand. From our launch video that has human actors, to our product details. I believe a distinctive, high quality, well implemented brand is still a hallmark of a strong product or service.
LLMs are so easy to swap out, so having good benchmarks/evals are pretty useful.
Even then, a lot of the time the model improvements are so obvious that you don't even need an eval.
What do you think that is referring to?