So I reverse engineered two dating apps
push32.com
push32.com
In the process revealing how erroneously so many put stock in matching via a system with a very clearly vulnerable system (compatibility rating algorithm) and went out of his way to optimize the system further with less than desirable results: he was the most desired Male profile and was going out on several dates a day and rejecting women he'd have coffee dates with in more and more absurd ways as he got more and more requests.
It was really funny, albeit cautionary, and I only caught like 65% of it, so if anyone finds it can you post it here?
Jack Dorsey's recent interview on the AI podcast hinted at how perilous this could be, too. The 'why' part of Machine learning/data Science/AI has to be just as tightly examined as the optimization development when deployed into Society.
And here's an article https://www.wired.com/2014/01/how-to-hack-okcupid/
I.e. boosting visibility probably brings proportionate benefits to the most desirable men, but little additional benefit to the ugly/short/older guy.
Edit: wtf @ stripping emoji HN
At least on OKCupid, women rate 80% of men as below-average attractiveness, while men rate women at right about 50% as below-average and 50% as above-average [0].
OKCupid has since deleted the post with their findings [1], but references to it still exist all over, as does an archived copy [2].
[0] https://techcrunch.com/2009/11/18/okcupid-inbox-attractive/
[1] https://theblog.okcupid.com/your-looks-and-your-inbox-8715c0...
[2] http://web.archive.org/web/20100324074028/http://blog.okcupi...
Why even ask for above/below average ratings? What people like/don'tlike is more interesting.
It's basically the same mechanism that works for physical attraction, i.e. if you had good interactions with people of a certain "type" then you'll be attracted to other people of that "type".
I looked at the links and didn’t see mention of those values, and I don’t know if they are commonly known, so I’m going to ask again.
Where are these numbers (95% of women and 50% of men) coming from?
https://medium.com/@worstonlinedater/tinder-experiments-ii-g...
https://qz.com/1051462/these-statistics-show-why-its-so-hard...
I can't find the actual statistic the above people are repeating, but these are similar.
That's it, I was really hoping to hear this in its entirety but could never find it: thanks!
It is nice marketing: look how easy it is to output entities with our framework!
The problem is: you almost never want to output complete entities. You want to output data based on the user's role/rights and on the context. I wish frameworks would put more focus on this.
That is, it's not just "/api/people/" that shows everything. You'd need: "/api/people/matched/" with a few fields, and /api/people/before-match/" etc. One end point per use case. Almost like one endpoint per... screen.
Suddenly, the old school MVC starts to make more sense. Your view dictates what data you need only for that view. Across resources. And it factors in entitlements. Often, people think this means no AJAX or SPA. That's not quite right, they are orthogonal.
What you really want is a data-on-the-wire protocol but with RPC calls that are still domain based rather than resource based. Everything should be its own endpoint. Django Forms and Formsets, and the Rails and Laravel equivalents are much closer to this than REST (even if they can be RESTful themselves). Think "UpdateProfile" instead of "UserResource". The UpdateProfile may have a lot more fields, some which don't even belong on the User database table.
Lot of rambling here, but I've been meaning to consolidate this in a blog post.
Some related thoughts here: https://news.ycombinator.com/item?id=21875331
I think we might already have all the architectural solutions we need, we just haven't put together how to express the architectures we want in the new technologies we have.
Some combination of the onion model/clean architecture, CQRS, APIs, data federation like GraphQL and microservices is probably how it can all go together. I know I am mixing implementations and patterns together in this sentence, but am just trying to express patterns and examples of things that have pieces of those patterns implemented.
Just apply that here - some API endpoints should be backend "system APIs" offering as much access to entities as possible - think "data layer APIs".
When you want to expose data to consumers for use in a UI, these are "experience APIs". These are APIs that might be built as part of the backend that directly interacts with the frontend.
Mulesoft talks about this a lot [1]. I am a user of it at work, not necessarily a big fan, but I think their architecture idea is applicable anywhere and helpful for thinking about layers of APIs in a large system. It's similar to the onion model, just applied to APIs.
Now, if you are building a smaller system this might all seem overkill. For a small system it might be true, unless you believe it will scale and need to be architected like this from the beginning. I think scale of the end result should always be factored into architecture decisions, because not one size fits all.
I think personally there is a model out there that hasn't been expressed fully/fully realized, that integrates ideas of APIs, GraphQL, CQRS, microservices and the onion model.
[1] https://blogs.mulesoft.com/dev/api-dev/what-is-api-led-conne...
Many of these technologies (GraphQL) seem like layers on top of layers. Really, I just want a page with 3 fields specific to that page. They may come from 3 different tables.
GraphQL's original solution to this was to batch the request to 3 different resources, but the underlying problem was still there. Of course, you can make ad-hoc GraphQL equivalents to a random REST ad-hoc /api/get-info-for-my-profile-page/.
The problem becomes especially noticeable when you want to update nested resources. Say, a profile's subscription status. Instead, in "MVC" world, I would just bump up the subscription status to be a first class field on that view. So no profile->subscription->status. Just "subscription_status". Hope that makes sense.
So far, I've reverted to using Django Forms (across tables) and breaking them down to JSON to use them in a SPA. My endpoints are still very much domain specific, but I can use XHR/fetch to make the page interactive.
I don't have much experience with such things. Can such an analysis be done with a non-rooted device?
Modern iOS/Android have something called SSL Pinning. Which means if you want to use a self-signed cert to Man in the Middle the HTTPS traffic it won't work. So you need to patch the network stack to allow this. I have a spare iPhone I keep rooted for this reason.
(I know things like Google Apps and such use it, but I'm not sure about less popular ones)
Rooting the device enables you to put custom certificates in system store, and bypass this check.
That said, it is more convenient to have a spare rooted phone.
Yep, the mobile OSes have that option, but it’s not very commonly used. I fairly regularly snoop traffic on apps and it’s not super often that I’m blocked temporarily by cert pinning.
I can already imagine the amount of technical barrier and knowledge gap one needs to fill even before getting started..
Holy shiet. It's impressive.
have to do a lot of multi-system level debugging, and/or low-level debugging and optimization
Time in front of the screen is all you need. All code translates to an execution layer.
It helps if you have a few years of dev under your belt. Bonus for low level languages like C or assembly.
Learning how this stuff works in the forward direction makes spotting patterns a whole lot easier. It’s a lot easier to start RE when you’re already familiar with stuff like calling conventions or memory layout (for example).
From there, there isn’t a ton of formal education as far as I’ve seen. I am really fond of Smash the Stack’s IO wargame if you’re interested in CTF-style challenges. I also spent a good bit of time compiling my own small programs and then using them to learn the tools. When you’re starting off, RE is a lot easier when you know what you’re looking for.
Fundamentally though, you will probably need to build up some low level knowledge of whatever you are targeting (whether it's an app platform, TCP packets, C/ASM code, etc).
If you have a web background there are lots of sites that have fun challenges you can start with. You could try Google's XSS game: https://xss-game.appspot.com/
It's a field that really interests me since I always liked puzzles and games and am happiest at work debugging tricky issues - but I'm not sure how easy it would be to break in as a 'junior' at age 35+.
I know that's vague but there are companies that make new products that for one reason or another need to talk to old products and the source, APIs, even binary interfaces are opaque, undocumented, obfuscated or otherwise unaccessible.
Developer tools, CAD/EDA, embedded systems... lots of people are building new stacks that have to talk to older ones at some point.
https://jobs-noblis.icims.com/jobs/8246/junior-reverse-engin...
https://jobs-noblis.icims.com/jobs/6640/software-reverse-eng...
https://jobs-noblis.icims.com/jobs/6415/software-reverse-eng...
But in the past I've worked with malware analysts using IDA Pro to reverse engineer stuff. They worked for SAIC. That's another place to try.
I've had a confidential security clearance for years because of my job working on government military applications. I remember the process being insanely rigorous. I wonder how much longer it would take to get a Top Secret clearance?
Other anecdata: 60% of my cohort got a clearance in less than 8 months (I believe their sponsoring project paid for expediting?), at least one person was >3 years in queue last I checked.
4) I would say is the best and the safest solution at really low cost. Followed by 1) at price/value ratio.
The problem is that it means the requests can't really be cached, so for static assets it's not ideal. S3 supports etags though but I've never touched that so I don't know how well it all works; I personally don't use signed URLs for stuff I want cacheable.
Another problem is if the random image id is leaked.
The other solutions you have where you authenticate the path as such, which should ideally be used for more sensitive assets:
- Proxy asset requests through application server, let it authenticate the user access. You can even generate short lived S3 urls this way, and your objects don't have to be public.
- Proxy asset requests through a reverse proxy like Nginx, authenticate there, you have a lot of options here like using something like JWT to authenticate the resource access.
- Proxy asset requests through your reverse proxy CDN like Cloudflare, authenticate there, like in case of Cloudflare JWT is supported.
Protecting from MITM is a separate (important) concern for that approach to be valid though, and it's not 100%.
If you use an integrated cloud+cdn solution, the problem is largely solved for you, assuming you pick up the best practices and run with them
Good reminder that fun little Easter eggs like using the bizarre 418 return code can bite you. I’d hate to be the engineer defending that decision after this vulnerability was discovered.
For example, I have all my static assets on S3 and I want to generate a link that will make data available for a long time (let's say 1 year) but with S3 signing you can only generate a link available for a week max.
How would you go about doing this without relying on another server?
I think this is what happened with the public bucket. They thought about how to deliver static assets without relying on a server and the only way they found is to make the bucket public.
If you want "clever" ideas, maybe a Lambda that moves off objects of a certain age to a private bucket?
I disagree on what the issue was with their S3 bucket. As these were all public static assets, the real problem was just the ability to bulk enumerate them. As mentioned in the post, the two issues were: 1. ListObjects was enabled 2. The filenames lacked sufficient entropy (debatable in my opinion)
So far I haven't seen any big drawbacks. It does mean storing the same objects multiple times in S3. But S3 storage is relatively cheap unless you have a huge amount of data. If bandwidth was ever a problem, it would be simple enough to wrap the transient prefix in a CDN.
Never made a blog post. Maybe I should.
Unfortunately most startups are obsessed with growth and don't really care about security, privacy, etc.
People can't all devote time and energy to everything that's important. It's too big a world. We shouldn't conclude that those things don't matter to them. Our current emergency is a fine example: until a pandemic happened, few knew enough to worry about it. But that doesn't mean we were indifferent to the outcome.
Well if the users can't spend any time on it, it isn't sufficiently important for the business to care about it as the users will be briefly miffed and then go on to what they consider important.
User unhappiness matters only so far as it changes behavior.
Industrial food companies were taking significant liberties with food quality and safety. And it was apparently fine! People appeared not to care. Then Upton Sinclair's book The Jungle happened to catch the popular imagination, leading to a wave of outrage and investigations. People still "didn't care" in the sense of, say, not buying industrially produced food. But they did very much care, leading to a regulatory regime that has lasted more than a century.
They show that expectation by withholding their money or not using the product when something bad happens. I don't see them doing much of that when companies have data breaches.
If these things caught the sparkle in customer's eyes, we'd all know it by now.
In terms of popular media, I thought Mr. Robot did a pretty good job of making privacy/security tech seem sexy, even if the main character is usually bypassing it.
Here’s a good example: https://youtu.be/i9CBKGLVCME
If we don't come up with something like that as an industry, eventually somebody else is going to do it for us. And we won't like that one bit.
Strongly disagree. It's just the simple reality that most users don't care about security. The vast majority of potential consumers in the world don't choose digital products based on security. I always see this security angle touted on Hacker News, but I'm quite frankly shocked that people here don't have the self-awareness to realize that we live in an uber-tech geek's echo chamber.
Have you ever met an "average" Facebook user? They really, truly, do not understand or care about security. I'm very confident that even if you sat one down and walked them through all of the implications of what poor security even means, they would walk away and not change their behavior whatsoever.
Adopting the stance that "vast majority of potential consumers in the world don't choose digital products based on security" time-and-time again bites organizations in the ass when there's a breach.
The bite isn't very hard though. The largest data breach of the 21st century in terms of users was Adobe and it cost them just 2 million in legal.
The only painful data breach I can think of financially has been Equifax. Everyone else just sent out a "reset your password" email, paid for a couple lawyers and PR people, and went on with their companies.
Can you name a company killed by a data breach? I can't think of one.
As it is what are the odds of any one user running into issues because of lack of security and privacy? It seems fairly low.
It seems like an opportunity for Apple (assuming they're actually better) to run some scare ads (99% of people scammed via there computer were running Windows/Android) if that's true. If it was true a good ad campaign could get people to care?
For internet stuff how about 98% of the people who got their bank accounts hacked were hacked by leaks of data on Facebook. (probably not true and not provable)
Maybe we need some security insurance who will then audit software and only insure customers that use certified software? They'd have an incentive for their audits to be good because they pay out if it turns out the software is not secure And if their market was big enough then software creators would want to be certified.
It's even possible some standard sandboxes could help make it easy to certify. Add this sandbox to your app and you're certified? Maybe some OSes that already have sandboxes would automatically get certified but server side you'd need audits?
Just throwing out ideas.
Amazon/MS/Google should do a better job of making it hard to leave things unprotected. They are no longer a new service and no longer have the excuse of having to avoid friction.
Users broadly have no way of analyzing security + privacy. Even most software developers don't have the time or expertise to reverse engineer and analyze even one app, let alone every release of every app they use. They just have to take it on trust. For things like cars we have mandatory standards for vehicle design, crash testing requirements, fuel efficiency and emissions standards etc. to try and make sure that people can expect a certain level of performance and safety. For software there's nothing like that.
* HackerOne is very expensive for most startups, IIRC it's around 30k USD/year only to be hosted on their platform (the price is probably different for different companies though.)
* If one doesn't want to use a platform like HackerOne you're in a tax hell doing payments to people all over the world
* The cost of paying for the bugs reported which can be massively expensive if you want to be competitive.
* The time it takes to handle all reports is massive. You might imagine you get cool and interesting research but 99.9% are people copy-pasting output from automated scanners, people reporting on out of scope things, usually low severity with things like "missing x-frame-options" or just straight up false positives. It's both frustrating, demotivating and time consuming.
I think a self hosted responsible disclosure program is better and more sustainable for most startups. Add a security.txt [0] to the website. You might still get a bunch of low quality reports but at least you give people a structured way of disclosing findings.
Why would the company paying be in tax hell? The recipient is the one who needs to pay income tax
I think it has do do with how bug bounties generally work. Since bug hunters aren't sending an invoice from a registered company you have to pay employer tax or similar. But as I mentioned, I'm not entirely sure since I don't work with that part. It might also differ depending on where the company is registered.
For the love of all that is holy... have your AWS, Azure, etc. cloud service setup/config wire-brushed by a 3rd party expert before deploying!
That, or perhaps these cloud services need a clippy like wizard: "Are you sure these files should be 777? Perhaps you want 644?".