239 karma · joined December 16, 2024
X/Twitter: @adidoit LinkedIn: https://www.linkedin.com/in/adipradhan/
Reach out hello@socratify.com
Most of the high volume enterprise use cases use their cloud providers (e.g., azure)
What we have here is mostly from smaller players. Good data but obviously a subset of the inference universe.
With better imaging, tooling, and archaeological funding, I'm sure we'll find much more evidence like this
So many countries bronze and ancient ages are underexplored
Getting $200 subscriptions from a small number of whales, $20 subscriptions from the average white-collar worker, and then supporting everything us through advertising seems like a solid revenue strategy
Anthropomorphism of LLMs is obviously flawed but remains the best way to actually build good Agents.
I do think this is one thing that will hold enterprise adoption back: can you really trust systems like these in production where the best control you can offer is that you're pleading with it to not do something?
Of course good engineering will build deterministic verification and scaffolds into prevent issues but it is a fundamental limitation of LLMs
The more prevalent automation is, the worse humans do when that automation is taken away. This will be true for learning now .
Ultimately the education system is stuck in a bind. Companies want AI-native workers, students want to work with AI, parents want their kids to be employable. Even if the system wants to ensure that students are taught how to learn and not just a specific curriculum, their stakeholders have to be on board.
I think we're shifting to a world where not only will elite status markers like working at places like McKinsey and Google be more valuable but also interview processes will be significantly lengthened because companies will be doing assessments themselves and not trusting credentials from an education system that's suffering from great inflation and automation
The overton window on LinkedIn is actually quite small and because everyone there is really an employee rather than an employer, you get essentially slop that has been easily trained on and therefore is easily generatable by AI. It's just all low perplexity takes.
There's mostly no room for nuance because of the performative takes. Unlike a forum like Hacker News where your identity is almost totally abstracted away, every LinkedIn post is a move in the status game of career visibility.
There are adjacancies in white collar work like financial analysis that they will go after. All these will capture high ARPU usage.
Consumer is not their only path to revenue but it is probably the easiest to model. The enterprise play to automate and accelerate some white collar workers is a clear target not reflected here.
It planned way better in a much more granular way and then execute it better. I can't tell if the model is actually better or if it's just planning with more discipline
It's simply too complex to fix. I think we'll see increased investment by corporates who do keep hiring on remediating the gaps in their workforce.
Most elite institutions will probably increase their efforts spent on interviewing including work trials. I think we're already seeing this with many of the elite institutions talking about judgment, emotional intelligence critical thinking as more important skills.
My worry is that hiring turns into a test of likeability rather than meritocracy (everyone is a personality hire when cognition is done by the machines)
Source: I'm trying to build a startup (Socratify) a bridge for upskilling from a flawed education system to the workforce for early stage professionals
Even imperfect assistants increase leverage.
From my perspective the challenges for vendors and SaaS providers are [1] discovery [2] monetization [3] disintermediation
I think it's less of a concern if you're Shopify or those large companies that have existing brand moats.
But if you're a startup, I don't think MCP as a channel is a clear-cut decision. Maybe you can get distribution but monetization is not defined.
Also I'm sure the model providers will capture usage data and could easily disintermediate you , especially if your startup is just a narrow set of prompts and a UX over a specific workflow.
The Reforge guys have been talking about a channel shift and this being it but until incentives are clear I'm not sure this is it yet. Maybe an evolution of this.
I'm building an AI coach for job seekers / early stage professionals (Socratify) and while I'd love more distribution from MCP UI integration I think at this point risk is higher than reward...
It's still early in the paradigm and most startups will fail but those that succeed will embed themselves in workflows.
it's art and meaning and entertainment all in one
if it's been a while you should read them again with fresh eyes.
Why would you not at least link it to the pro and ultra accounts
at least you could upsell the pro subs to ultra. Millions of claude code and codex users who are into agentic coding is your servicable market paying attention today.
Now I'll delete antigravity and go back to codex / claude code / cursor ...
Claude code seems to be more compatible with the model (or the reverse) whereas gemini-cli still feels a bit awkward (as of 2.5 Pro). I'm hoping its better with 3.0!
The word "spec" is a bit overloaded and I think we're all using it to define many things. There's a high-level spec and there are detailed component-level specs all of which kind of co-exist.
However short and targeted specifications at the right level of detail and fidelity, can be extremely useful during coding with agents
It's a bit weird it took Anthropic so long considering it's been ages since OpenAI and Google did it I know you could do it through tool calling but that always just seemed like a bit of a hack to me
It's a form of enshittification perhaps. I personally prefer some of the GPT-5 responses compared to GPT-5.1. But I can see how many people prefer the "warmth" and cloying nature of a few of the responses.
In some sense personality is actually a UX differentiator. This is one way to differentiate if you're a start-up. Though of course OpenAI and the rest will offer several dials to tune the personality.
Career Skills AI Coach. Sharpen how you think and speak by debating AI
We are clearly on the verge of the largest white-collar skills dislocation ever. Our goal at Socratify is to make skill building and reskilling for interviewing, onboarding, promotions, and career change as effective as possible with an AI coach and sparring partner.
AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of how well agents really work. We introduce the Holistic Agent Leaderboard (HAL) to address these challenges. We make three main contributions. First, we provide a standardized evaluation harness that orchestrates parallel evaluations across hundreds of VMs, reducing evaluation time from weeks to hours while eliminating common implementation bugs. Second, we conduct three-dimensional analysis spanning models, scaffolds, and benchmarks. We validate the harness by conducting 21,730 agent rollouts across 9 models and 9 benchmarks in coding, web navigation, science, and customer service with a total cost of about $40,000. Our analysis reveals surprising insights, such as higher reasoning effort reducing accuracy in the majority of runs. Third, we use LLM-aided log inspection to uncover previously unreported behaviors, such as searching for the benchmark on HuggingFace instead of solving a task, or misusing credit cards in flight booking tasks. We share all agent logs, comprising 2.5B tokens of language model calls, to incentivize further research into agent behavior. By standardizing how the field evaluates agents and addressing common pitfalls in agent evaluation, we hope to shift the focus from agents that ace benchmarks to agents that work reliably in the real world.
The use case was to build a knowledge graph to drive recommendations for the next best thing the user should learn.
After a few weeks of getting frustrated I went back to good old Postgres and writing a few tools for agentic retrieval.
It seems the agents are smart enough to traverse a database in a graph like manner if you provide them with the right tooling and context
https://www.olpejetaconservancy.org/what-we-do/conservation/...