All the data can be yours: reverse engineering APIs
jero.zone
jero.zone
I learned a lot from that experience, and it's also just plain fun to do. We did get some strongly worded letters from Robinhood though, lol. They tried blocking our servers but we just set up this automated system in Digital Ocean that would spin up a new droplet each time we detected a blockage, and they were never able to stop us after that.
Fun times.
Sadly the CR forums are gone, so my rather popular thread that had feedback and support is long gone.
Like sure then you can add in hertzer or w/e and keep adjusting but idk if somebody keeps ban dodging by using the same provider it seems like you'd just try banning that provider early on?
(Hmmm, all the devices we used then would have just stopped working with the shutdown on the 3G network here. I wonder if it's all broken, or if they've upgraded all those devices to 4/5G ones?)
Unsurprisingly, the most sleazy players are the first ones to go after someone accessing their services in ways they didn't anticipate or intend :).
Our CEO was friendly with an investor who had an account at some big Singaporean trading firm called Lim Tan. He gave us his account credentials and I began working. A few days later my boss comes over to my desk and says "stop whatever you're doing right now." Apparently my traffic had set off so many alarm bells that the CTO of Lim Tan was woken up at 2am. They permanently banned the investor, which I felt really bad about. What's crazy is that I wasn't even doing anything weird. I was just poking a bit at their authentication methods. That was when I learned that Singapore tech doesn't fuck around.
Faking real user clients won't prevent these alarms.
[1] https://nhl-remix.vercel.app/ [2] https://ahl-remix.vercel.app/ [3] https://pwhl-remix.vercel.app/
import asyncio
from kickerde_api_client import Api
from kickerde_api_client.model import LeagueId
api = Api()
query = {'league': LeagueId.NFL, 'season': '2024/25'}
season = asyncio.run(api.league_season(**query))
print(season['longName']) # 'National Football League'
print(season['country']['longName']) # 'USA'
print([
team['shortName']
for team in season['teams'].values()
if team['shortName'].startswith('B')
]) # ['Buffalo', 'Baltimore']
day = season['gamedays'][18]
print(str(day['dateFrom'].date())) # '2025-01-05'
[0]: https://kickerde-api-client.readthedocs.io/en/stable/autoapi...I gave him the perfect solution: Why don't you just call them and ask how much money they want for access to their product catalog? Give them tiny bits of information if they don't approve immediately. Tell them it is for one of their customers and that their site is to slow. They are the only supplier not in the interface which is bad for them and bad for you. If they still refuse offer them access to their competitors data at a modest fee.
He couldn't stop laughing, he never considered it. When the sun set the next morning he put in the call. They gave him access immediately, they were so happy to finally get rid of his crawler messing up their analytics data. Apparently the people on the other end of the tube also didn't sleep. They had a good laugh about it.
I'm on the dashboards and integrations team, and I don't have direct access to the codebase of the main product. As the internal APIs have no documentation at all, I'm always "hacking" our own system using the browser inspector to find out how our endpoints work.
I've seen a case where one team developed a temporary hack to automate some process in their product, then allowed another team making a sibling product to use it to test a possible feature; soon after, the possible feature became an actual one, everyone seemingly forgot the API was a throwaway test. Over the years, as both products evolved, that feature got pretty flaky, and in one of many cross-team attempts at fixing it, someone from the original time finally pointed out that the whole thing is still relying on a temporary hack in the original product that was never intended to be productized...
That's a common story; in my university (in Padova, UniPD) happened something even worse. They tried hard to shut down an unofficial app (Uniweb) that was installed by most of the students in favor of the "official" one, that was completely unusable (and probably was born out of a rigged contract). At the end the best one won and became official, but that was after a lot of struggle.
Having said that, speaking as a Macalester alumnus, I wouldn’t put it past them to be a bit petty behind the scenes :)
Sounds like what happened with the Apollo app and Reddit
Reverse engineered just enough of that .net monstrosity to get the data I cared about. Replacing the school's portal with my own script made me so happy. They kept breaking my scraper constantly though, got to the point it was too much effort to maintain it. Worst part is I don't even think they were doing it on purpose. They were just bad in general and couldn't keep their site together.
Its especially annoying since many use binary message formats and there isnt a great way to document an arbitrary binary message protocol.
A couple techniques im trying out:
- websocat and wsrepl for reverse engineering interactively: https://github.com/doyensec/wsrepl
- kaitai struct for documenting the binary message formats: https://kaitai.io/
After that I used ImHex, pretty much exactly like in that blog to reverse engineer the websocket packets. The DSL is a little finicky but once you wrap your head around it, it`s very nice and powerful.
I don't dare monetize that project, but I wish I could use my skills to make some extra money at reversing APIs. Wouldn't know where to begin though.
The game was a multiplayer FPS and involved cosmetic micro transactions.
Due to mismanagement they shut down in one year. It launched in late 2022 and EOSed late 2024.
I'm in the US, yes. The company is Japanese though and they may not care as much about US law when choosing to pursue. Even if they ultimately don't have a case, there's still enough gray area to torment-by-lawyer.
For that reason I chose not to take any cash for the project. I also don't distribute any game files. Trying to minimize risk as much as possible.
But yeah! It was a good community building project. We're about to hit 10,000 members I hope it leads to some connections that are profit rearing.
Whatever you're doing, I hope you get away with it.
Also, I made a typo. They shit down in late 2023, not 2024. So between November 2023 and March 2024, there was no public place to play the game we've brought it back since though.
Hope Bamco lets us be in our little corner for now!
It's just an endless cat & mouse game.
Getting into the internals of an LMS is tricky, but surprisingly very interesting and very fun.
It's much more straightforward if you can find a GraphQL, Swagger, or OpenAPI spec to automate conversion I'd imagine.
Glad to see this being called out. Sure, I get why it's convenient. Misspelling a field by one character is a daily occurrence ("activty" and "heirarchy" are my regulars). The catch is that spellchecking queries and returning valid fields in the error effectively reduces entropy by both character space and message length, varying by the type of distance used in the spellcheck.
https://gizmodo.com/ring-s-hidden-data-let-us-map-amazons-sp...
Teams that I've seen working on apps now implement much stronger checks on APIs especially Android apps such as SafetyCheck and DeviceCheck and other methods, which makes using strings rather basic to see them.
And most apps are now encrypted so you just see junk in the logs.
Granted, smaller fish like the ones OP is referring to generally don't have aggressive anti automation measures in place, so it can be easy...but generally these techniques don't work if the operator has put the proper measures in place.
lmao ok
Not that I'm hoping for it, I too like to play around like OP. But I'm surprised how little I've encountered it in the wild.
[0]: https://github.com/search?q=org%3Aauth0%20DCDevice&type=code
I also implemented a bot mitigation system for a large international company so I got to see techniques used from the other side. Mobile phone farms and traffic from China was the most difficult to mitigate.
YouTube is a video monopolist - they know that it is an awful service for viewers, but there's no need to improve things because where else would viewers go?
Unfortunately, they are also part of a conglomerate that has reached full dominance in several online markets that are not growing. And they need to show quarter-on-quarter growth to shareholders. So you can't capture more market share, and there aren't any more people discovering the Web, so you stick additional ads everywhere. These people know exactly what they are doing.
So I setup a Raspberry Pi with the SSID as the camera, logged the first API call. Then I ran that API against the camera, and then wrote a small python script on the Pi to return the same API result. I did that call by call until I'd worked everything out. I wrote it up here: https://www.hotelexistence.ca/reverse-engineer-akaso-ek7000/
Are there any good tutorials for that? 'strings' is not the greatest name for searching good information.
https://www.corellium.com/blog/ios-mobile-reverse-engineerin...
(No affiliation.)
Edit: After some Google-ing and trying it on my macbook, there is a native CLI tool called "strings". Supposedly it does the following: strings is primarily used to find and display printable character sequences in binary files, object files, and executables. Which means the author is probably looking at the app to see the hardcorded characters in the app binary(?) and searching for the API end points.
strings file
will tell you all of the ASCII strings in file. strings -el
For utf-16 encoded ascii strings (very useful when dealing with windows executables).While most of the time, you're dealing with variables and such in programs, at some point you have to hardcode some information such as URLs to query so something like
BASE_URL = "https://example.com" result = requests.get(BASE_URL + "/api/blah"
If we pretend this is in an Android app which is stored as an apk file (a zip file basically), running strings would spit out "https://example.com" and "/api/blah"
It'll also spit out anything that appears to be an ASCII character so plenty of junk but it's often quite handy as a starting point.
There are, of course, much more precise tools such as man in the middle proxying but that you'll only capture traffic for endpoints actually used by said app. The app may contain other endpoints let unused, rarely triggered and so on.
https://httptoolkit.com/blog/frida-certificate-pinning/
https://github.com/httptoolkit/frida-interception-and-unpinn...
I agree with sibling commenters that automating the browser with DOM APIs is an easier route to go though.
I deleted a tweet and saw this request:
HTTP POST https://x.com/i/api/graphql/VstuveVgh5q5jk7lmnVopqr/DeleteTweet
{
"variables": {
"tweet_id":"12344567899123",
"dark_request":false
},
"queryId":"VstuveVgh5q5jk7lmnVopqr"
}
You can execute these from javascript in the browser if the auth part is too complicated.### Update, this is the pure javascript console way, if you don't want to write your own client doing HTTP posts
I played with the console more and got these parts:
// Find all tweets on screen (this gives you the tweet IDs too)
document.querySelectorAll('a > time')
// Click the "more" button on the first tweet document.querySelectorAll('a > time')[0].parentElement.parentElement.parentElement.parentElement.parentElement.parentElement.parentElement.parentElement.parentElement.parentElement.parentElement.querySelector('button').click()
// Click delete on the tweet document.querySelectorAll('[data-testid="Dropdown"]')[0].children[0].click()
// Confirm delete document.querySelectorAll('[data-testid="confirmationSheetConfirm"]')[0].click()You could do something similar; as someone else suggested, just walk the feed via DOM elements.
https://thecopenhagenbook.com/
It was shared last or week or so on here.
What's the best channel to reach you? I'll email you if you are not seeing this comment.
Give ‘em a shout out by saying “hi” in a request parameter or something while you’re reverse engineering.
Instead, I'd recommend reaching out to the website owners directly to discuss your API needs - they're often interested in hearing about potential integrations and use cases.
If you don't receive a response, proceeding with unauthorized API usage is mostly abusive and poor internet citizenship.
Where I'll agree with you is cases where people do this and impose way more traffic than is typical and often more than is necessary (i.e. no caching). But that's not really specific to reverse engineering apis, that's just about being a good internet citizen in general.
I'm of the opinion that a user agent is a user agent, and website owners shouldn't pick and choose what user agents they support. Target the behaviors that affect your infrastructure, not the means of access and algorithms used to process what you send.
I agree but that's not usually how it goes. From what I've seen, it's mostly very poorly written scripts with not rate limiting and no backoff strategy that will be hitting your api servers.
Casting shade on API reverse engineering when what you actually have is a failure to rate limit is throwing the baby out with the bathwater. Abusive users will abuse until you build in a technological method to stop them, and user-agent sniffing provably doesn't work to stop bad actors.
The concept of a flexible, customizable User Agent that operates on my behalf is a key idea that's foundational to the web, and I'm not willing to cede that cultural ground in the vague hope that we can make the bad guys feel bad and start just using Chrome like civilized people.
Bots looking for exploits is rude, spamming an endpoint with more traffic than normal is rude.. but a human trying to figure out the API that you exposed to the internet? That's just fair play.
Also, better to ask for forgiveness than to ask for permission. The author is adding value to the world while hurting nobody, and the answer would likely be an automatic "no" anyway.
If you expose an API, and you want to tell a user that they are "unauthorized" to use it, it should return a 401 status code so that the caller knows they're unauthorized.
If you can't do that because their traffic looks like normal usage of the API by your web app, then I question why their usage is problematic for you.
At the end of the day, you don't get to control what 'browser' the user uses to interact with your service. Sure, it might be Chrome, but it just as easily might be Firefox, or Lynx, or something the user built from scratch, or someone manually typing out HTTP requests in netcat, or, in this case, someone building a custom client for your specific service.
If you host a web server, it's on you to remember that and design accordingly, not on the user to limit how they use your service.
Just because you're right doesn't mean you aren't wrong.
Their very reasonable question was: if you can't distinguish the reverse engineered traffic from the traffic through your own app in order to block it, then what harm is the traffic doing? Presumably it's flying under your rate limits, and the traffic has a valid session token from a real customer. If you're unable to single it out and return a 4xx, why does it matter where it's coming from?
I can think of a few reasons it might, but I'm not particularly sympathetic to them. They generally boil down to "I won't be able to use my app to manipulate the user into taking actions they'd otherwise not take."
I'd be interested to hear if there are better reasons.
If you really believe this you'll use a custom user agent instead of spoofing Chrome. :-)
Some websites use HTTP referer to block traffic. Ask yourself if any reverse engineer would be stopped by what is obviously the website telling you not to access an endpoint.
I'll add that end users don't have complete information about the website. They can't know how many resources a website has to deal to reverse engineering (webmasters can't just play cat and mouse with you just because you're wasting their money) nor do they know the cost of an endpoint. I mean, most tech inclined use ad blockers when it's obvious 90% of the websites pay the cost of their endpoints by showing ads, so I doubt they would respect anything more subtle than that.
That endpoint will be expensive regardless of whether it's your own app or a third party that's calling it too often, so design it with that in mind.
Your app isn't special, it's just another client. Treat it that way.
If you could ensure that the web server can only be accessed by your client, you would do that, but there is no way to do this that can't be reverse-engineered.
Essentially your argument is that just because a door is open that means you're allowed to enter inside, and I don't believe that makes any sense.
Imagine trying to validate that all letters sent to your company are written by special company-provided typewriters and you would run into the same fundamental limits.
Whenever you design any client/server architecture, the first rule should always be "never trust the client," for that very reason.
Rather than trying to work around that rule, put your effort into ensuring that the system is correct and resilient even in the face of malicious clients.
Read up on the history of User Agent string, and why everyone claims they're Mozilla and "like Gecko". Yes, it's because of all the silly people who, since earliest days of the WWW, tried to change what they serve based on the contents of User-Agent header.
https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim...
https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim...
(In the United States at least)
> Q: I thought that United States currency was legal tender for all debts. Some businesses or governmental agencies say that they will only accept checks, money orders or credit cards as payment, and others will only accept currency notes in denominations of $20 or smaller. Isn't this illegal?
> A: The pertinent portion of law that applies to your question is the Coinage Act of 1965, specifically Section 31 U.S.C. 5103, entitled "Legal tender," which states: "United States coins and currency (including Federal reserve notes and circulating notes of Federal reserve banks and national banks) are legal tender for all debts, public charges, taxes, and dues."
> This statute means that all United States money as identified above are a valid and legal offer of payment for debts when tendered to a creditor. There is, however, no Federal statute mandating that a private business, a person or an organization must accept currency or coins as for payment for goods and/or services. Private businesses are free to develop their own policies on whether or not to accept cash unless there is a State law which says otherwise. For example, a bus line may prohibit payment of fares in pennies or dollar bills. In addition, movie theaters, convenience stores and gas stations may refuse to accept large denomination currency (usually notes above $20) as a matter of policy.
> In short, when a debt has been incurred by one party to another, and the parties have agreed that cash is to be the medium of exchange, then legal tender must be accepted if it is proffered in satisfaction of that debt.
You are correct that if cash is not accepted at all, or if payment is to happen ahead of the exchange of goods or services, you are not obligated to accept arbitrary cash.
And I never claimed otherwise
In other words, an agreement isn't required in order to refuse legal tender, an agreement would be required to make it mandatory.
A court might decide that an agreement to accept cash without specifying in what form was meant to include dimes, but I see no evidence anywhere that a court has to rule that way if the contextual evidence suggests something else was probably meant.
The law says that coinage is valid legal tender for an offer to settle a debt but the counterparty is not required to accept it... unless they contractually agreed to do so.
Only the U.S. government is required to accept payment in coins. Many states also require their agencies to accept payment in coinage but some have laws limiting the size of debts that can be paid this way.
Think logically about this. What do you think legal tender even means otherwise? Why would you need a special term to denote a form of payment that a creditor can accept if they want to? I could accept settlement in jelly beans if I wanted to. The entire point is that you must accept legal tender, that is what makes it different from everything else.
Sometimes you can get the green light just by reading API docs/School's privacy policy as they're usually obliged to have one (ofc this primarily applies to school APIs like in OP's article)
I’m not sure taking liberties with things you don’t own, is always the best policy, nor is putting the entire responsibility on the owner.
I don’t think this is something you can boil down to a simple black and white.
If you put a server on the public internet and I send it a message (assuming I’m not using ill-gotten credentials, etc.), anything it responds with is your problem, not mine.
No. It could be said, but wouldn't be true - objects in a domestic front yard are nothing like pamphlets placed on the sidewalk.
Not sure why you are acting so obtuse.
I am aware that lots of companies have ideas(TM) about how you should be able to use their products(TM) and may even add these to their Terms of Service, a document that has somehow become the last refuge for the bureaucratic organisation desperate to maintain control when forced to connect things to great unbureaucratic internet.
To that, I say: too bad. I never signed up for the new version of the internet and I do not consider TOS to be anything but noise. I used Pidgin back in the day and would again if it worked.
This absurd idea that website owners should have any say about what runs on your computer/device is nonsense.
The Times is really not a great tech company in any sense. If I were a bit less lazy/busy I'd get more into their audio app, but frankly their reporting has gone downhill. I guess they're running the referral mill strategy now with all the ads they put into the app where there were none. Maybe they can hire some better programmers, or better reporters, for that matter.
No, they don't get a say about what software you run on your computer. But if your computer is accessing private APIs that I pay for, then I get a say in how you get to use it. It's also up to me to secure the APIs and prevent abuse. If I don't do that then you're essentially free to do what you like with the API until such time that I do lock it down. I'm also free to block your IP address and delete your account if you break the rules of use of the API that I am paying for. Don't like it? Too bad. You can pay for infrastructure to run your own damn APIs.
For public APIs, the same rules about public usage of any physical space should apply. If you can see it "from public" aka logged-out, then you can take photos or record it (aka access the API). If it's a restricted area, then the public isn't allowed there and it's up to the entity trying to protect it to secure it.
I make my living for the last 7 years reverse-engineering non-public APIs from a service my company pays for. The service gets to set a rate-limit, and they enforce it. They know what we're doing and we are in contact often with their managers and engineers. They let us know if we're straining their systems and we respond by limiting use of some of their more expensive APIs. We've almost DDOS their system before, and this is a system millions of people subscribe to, that serves billions of pages per day. It's in everyone's best interest to get along and not abuse the APIs, and not cut us off from using them in a different way than they intended.
I would love it if this service took developers seriously and actually had a real developer program, but they do not, and they likely never will. It's more geared to consumers. But we depend on them in a very big way, so my job is reverse-engineering and scale up something that was never meant to be scaled. It's interesting work, but it also requires having an adult attitude and playing nicely with others. A little mutual respect can go a long way.
Adversarial interoperability is a cornerstone of the value of the internet and something we should fight hard to keep.