But I'd like to know an actual answer to this, too, especially since large parts of this post read as if they were written by an LLM.
They claim to have found that.
Which would be a major security hole. And sure, lots of startups have major security holes, but not enough that he could come up with these BS statistics.
I'm a little dismayed at how high up this has been voted given the data is guaranteed to be made up.
> Which would be a major security hole.
An officially supported security hole
https://platform.openai.com/docs/api-reference/realtime-sess...
Is Teja Kusireddy a real person? Or is this maybe just an experiment from some AI company (or other actor) to see how far they can push it? A Google search by that name doesn't find anything not related to the article.
The article should be flagged. Otoh, this should get discussed.
To be able to call the OpenAI directly from the front end, you'd need to include the OpenAI key, which would be a huge security hole. I don't doubt that many of these companies are just wrappers around the big LLM providers, but they'd be calling the APIs from their backend where nothing should be interceptable. And sure, I believe a few of them are dumb enough to call OpenAI from the frontend, but that would be a minority.
This whole thing smells fishy, and I call BS unless the author provides more details about how he intercepted the calls.
You mean, except for explaining what he's doing 4-5 times? He was literally repeating himself restating it. Half the article is about the various indicators he used. THERE'S EXAMPLES OF THEM.
There's this bit:
> Monitored their network traffic for 60-second sessions
> Decompiled and analyzed their JavaScript bundles
Also there's this whole explanation:
> The giveaways when I monitored outbound traffic:
> Requests to api.openai.com every time a user interacted with their "AI"
> Request headers containing OpenAI-Organization identifiers
> Response times matching OpenAI’s API latency patterns (150–400ms for most queries)
> Token usage patterns identical to GPT-4’s pricing tiers
> Characteristic exponential backoff on rate limits (OpenAI’s signature pattern)
Also there's these bits:
> The Methodology (Free on GitHub next week):
> - The complete scraping infrastructure
> - API fingerprinting techniques
> - Response time patterns for every major AI AP
One time he even repeats himself by stating what he's doing as playwright pseudocode, in case plain English isn't enough.
This was also really funny:
> One company’s “revolutionary natural language understanding engine” was literally this: [clientside code with prompt + direct openai API call].
And there's also this bit at the end of the article:
> The truth is just an F12 away.
There's more because LITERALLY HALF THE ARTICLE IS HIM DOING THE THING YOU COMPLAIN HE DIDN'T DO.
In case it's still not clear, he was capturing local traffic while automating with playwright as well as analyzing clientside JS.
How can he monitor what's going on between a startup's backend and OpenAI's server?
> The truth is just an F12 away
That's just not how this works. You can see the network traffic between your browser and some service. In 12 cases that was OpenAI or similar. Fine. But that's not 73%. What about the rest? He literally has a diagram claiming that the startups contact an LLM service behind the scenes. That's what's not described, how does he measure that?
You are not bothered by the only sign that the author even exist is this one article and the previous one? Together with the claim to be a startup founder? Anybody can claim that. It doesn't automatically provide credibility.
Presumably OpenAI didn't add that for fun, either, so there must be non-zero demand for it.
But I still believe the vast majority of startups do wrapping in their own backend. Yes, I read what he's doing, and he's still only able to analyze client-side traffic, which means his overall claims of "73%" are complete and total bullshit. It is simply impossible to conclude what he's concluding without having access to backend network traces.
EDIT: This especially doesn't make sense because the specific sequence diagram in this article shows the wrapping happening in "Startup Backend", and again, it would be impossible for him to monitor that network traffic. This entire article is made-up LLM slop.
He is not claiming to be doing that. He says what and how he's capturing multiple times. He says he's capturing what's happening in browser sessions. Reflect on what else you may to re-evaluate or discard if you misunderstood this.
> That's just not how this works. You can see the network traffic between your browser and some service.
Yes, the author is well aware of that as are presumably most readers. However for example if your client makes POST requests to the startup's backend like startup.com/api/make-request-to-chatgpt and the payload is {systemPrompt: "...", userPrompt: "..."}, not much guessing as to what is going on is necessary.
> You are not bothered by the only sign that the author even exist is this one article and the previous one?
Moving goalposts. He may or not be full of shit. Guess we'll see if/when we see the receipts he promised to put on GitHub.
What actually bothers is the lack of general reading comprehension being displayed in this thread.
> Together with the claim to be a startup founder? Anybody can claim that.
What? Anybody can be a startup founder today. Crazy claim. Also... what?
> It doesn't automatically provide credibility.
Almost nobody in this space has credibility. That could turn out to be Sam Altman's alias and I'd probably trust it even less.
In any case evaluating whether or not a text is credible should preferably happen after one has understood what was written.
He literally is. Look at the sequence diagram in the article, which specifically labels "Startup Backend". The whole article is made-up LLM slop.
This also matches the latency of a large number of DB queries and non-OpenAI LLM inference requests.
>Token usage patterns identical to GPT-4’s pricing tiers
What? Yes this totally smells real.
He also mentions backoff patterns, which I'm not sure how he'd disambiguate extremely standard backoff in a normal API.
Given the ridiculousness of these claims, I believe there's a reason he didn't include the fingerprinting methodology in this article.
Just because I'm frustrated with someone's inability to understand a text does not imply I want to defend or even personally believe what was written.
https://medium.com/@teja.kusireddy23/i-reverse-engineered-20...
The article is basically a description of where to look for clues. Perhaps they've contracted with some of these companies and don't want to break some NDA by naming them, but still know a lot about how they work.
This makes literally no sense. Why would any companies (let alone most of them) contract with this guy who seems hell bent on exposing them all.
The article is simple made up, most likely by an LLM.