I accidentally built a nudity/porn platform
elazzabi.com
elazzabi.com
- anything that allows anonymous file upload -> childporn + all of the above.
- anything that allows communications -> spam, harassment, bots
- anything that measures something -> destruction of that something (for instance, google, the links between pages)
- any platform where the creator did not think long and hard about how it might be abused -> all of the abuse that wasn't dealt with beforehand.
- anything that isn't secured -> all of the above.
Going through a risk analysis exercise and detecting the abuse potential of whatever you are trying to build prior to launching it can go a long way towards ensuring that doesn't happen. Reacting very swiftly to any 'off label' uses for what you've built and shutting down categorically any form of abuse and you might even keep it alive. React too slow and before you know it your real users are drowned out by the trash.
It's sad, but that's the state of affairs on the web as we have it today.
Put any technical book title into google with site:github.com in it and click the first PDF that shows up.
For fun, see https://github.com/topics/pornography
It's always fun to say that technically Pulp Fiction (released 1994, post-acquisition) is a "Disney" film.
On HN last year: https://news.ycombinator.com/item?id=20624074
A very good read!
Of course, points for creativity if they can still make lewd comments like "time for thrust attacks but whole" near a corpse bent over a banister.
The funny thing is that the game has incredibly sensitive username filter as well. I've fought countless players with names like "The #### K###ht" instead of "The ugly knight".
Doesn't Switch have ingame audio-chat at all?
Ugh.
Hasn't this been the meta since ever since?
Imagining you'd know all about that with camarades. It was my first exposure to web streaming as a teen way back in the day. Initially, it didn't strike me as a porn platform, but it didn't take long before it felt that way.
Content piracy to a much lesser degree as bitrates were so slow and expensive (though texts were fair game), but stalking, harassment, propaganda, bigotry, fraud, spam, etc., yes.
(FidoNet existed, but it was not easy to perform a full-content search of it, compared to eg. mail-archive.com now.)
The problem isn’t full-text search of a forum you’ve already signed up for.
Legend has it that if you include "Kibo" in your post to any of Usenet's more than 5,000 discussion groups, the all-seeing Kibo will respond to you.... [James "Kibo"] Parry spends about two hours per day as Kibo - tending a filter that collects any mention of "kibo" in any newsgroup.
https://web.archive.org/web/19961128094925/https://www.wired...
Tell us more!
related, SFW: https://virtuallyfun.com/wordpress/wp-content/uploads/2010/0... and https://de.wikipedia.org/wiki/ASCII-Art#/media/Datei:Schreib...
At one point there was a movie shot at SAIL with rather open minded approaches to computer-human interaction. In the early 1970's (as in the twenties?), "hangups" were for inhibited squares. (If they'd been running unix instead of waits, they'd probably have invoked nohup(1) for the shoot.)
Going back yet another century prior, we find the Victorian equivalent to OnlyFans: https://news.ycombinator.com/item?id=23791112
Camarades went 'downhill' about 6 months in when it went mainstream and more and more people that had other ideas of where we should go with it joined. Then they brought their buddies and it was 'game over', we had to decide what to do about it. This roughly coincided with the ad market collapse and drove the decision to put the porn behind a paywall to finance the rest of it. Which worked well for more than a decade.
I would create a company to have legal seperation for my private assets, i would use a ton of mechnism upfront to make sure i do my best to not support childporn and stuff and when i have analyzed how much work the proper way is, i will just stop thinking about it :)
And of course the normal stuff, links all prefixed with "nofollow", only whitelisted markdown content / HTML, &c.
The first year we had a handful of spammers try to post stuff, but it didn't take much time at all to filter it out. This year we didn't have any spam accounts at all.
So my suspicion is that getting ahead of the "spam problem" by doing heavy manual moderation early on is an investment. If he'd just spent 5 minutes a day looking at the accounts and deleting rubbish, he would never have been swarmed by spammers, and the moderation load would have remained relatively low, until he got popular enough that he could afford to actually spend some real effort automating / crowdsourcing the spam-fighting capabilities.
The downside, of course, is that they takes effort to maintain, and is a barrier to entry for new accounts.
I looked at nodebb might give that ago next
I wonder if you treated new content from unverified users like a temporary shadowban (they see it, nobody else does) how that would affect behavior.
Edit: Why does shadow banning feel like such an elegant solution? Everything has tradeoffs and I feel like shadow banning has tons of upside and very little downside. What am I missing?
Shadow banning only works when only authenticated users can see any content at all; then we arrange for only the offender to see content they have created. This works as long as the offender doesn't create multiple accounts.
(You need something slightly more clever, like allowing the content to be viewed from the same IP address as the last known address for that offender, plus some surrounding range (e.g. IPv4 class C subnet)).
Various soft fail mechanisms, including shadowbanning, degraded performance, errors, authentication failures, etc., may help avoid this.
The question is what adversarial model you're facing: is it some random pr0n / SEO / affiliate / fraudster, or is it a "nice website you gots heyah, be a shame if anyting happen' to it" squad?
Former can be modded away. Latter takes nuance.
One point of manual moderation is that spammers know it is happening.
We're always looking for little things that set our software apart!
As for your question: no, with the appropriate messaging, we don't often have repeated submission attempts.
I don't think it'd be an issue for a hardware forum, although mass approving comments and finding that one is defamatory about person or company X could get some lawyer's drool glands working for example..
Also section 230 appears to be under threat by the current administration and how things are interpreted - and I just read where Biden has been talking about removing it completely..
So this situation is in a state of flux at the moment I believe.
/not a lawyer / doctor yada yada
It does not care about what is moderated out. No one is responsible for speech that isn’t said.
I believe this would be more of an issue for some kinds of sites and much less an issue for others. (I doubt tom's hardware has ever had anything defamatory or libelous in posts to worry about ) - other places that moderate may do so in ways that are over-zealous in protecting the publishers and therefore that would create a different issue, but that's a different discussion.
Anyhow if you are referring to an anti-censor group or some kind of misinformation army like the bots and confangled news stories around the net neutrality debate, I would like to learn more about this/them and see why / how etc. thanks
Then you have moderating where people can post about anything - but when you get a complaint you research and take it down.
There are other forms and types of moderation between these.. which is why I said "in this way could open"
if you are moderating each message before if goes out, you are not acting so much as a platform as more of a publisher.. you could claim you did not know that an ad you published was not legal.. say it was turtles for sale, or someone offering weed for sale with 'codewords'.. you could claim that you did not know the girl in the ad not in a sex tape, that you just published the info...
But from what I have seen, if you allow everything to publish and take down stuff when notified you have a difference situation in regard to section 230, and claims in general, as compared to having to defend that you ad person could of should of known before publishing..
Indeed - section 230 doesn't care so much about what is moderated out in some ways, but the fact you are moderating out 8 year old's dating profiles, but not 13 year olds could become an issue- and the fact that you moderated out the 8 year olds in the first place could cause holes in certain defenses.
So you would not be responsible for that isn't said - but the fact you allowed one defamatory thing to publish while hiding a defamatory post about someone else - what isn't said could be an issue to prove that you have liability for what is.
Hope I'm saying that right, again this is my understanding of some issues that I have been involved in, and studies other people's issues and listened to their counsel - but I am not a lawyer.. if you have one and you are doing this kind of moderating perhaps have them look into the backpage cases and see how they pierced the "we're a platform, not an editorial publisher' thing. for one example.
However every so often you would have to create a new account. To do that the chmod had to be undone, but only for a minute or two. In that minute or two typically 2 or 3 spam accounts where created, and maybe 20 wiki pages spammed or defaced.
In short: defending the site had no effect whatsoever for us. The bots were always probing, every few seconds, and it never stopped even years after the site was totally locked down.
PS: SPAM wise the bots were just an annoyance. But every page hit runs moinmoin's Python code, and it's not the fastest thing. We were running on low end VPS's that took a dim view of anybody using too much CPU. Our VM regularly got shut down because of those bloody bots.
Right, so I think the key think spammers need to make spam "pay" is to make it scalable: only 0.1% of people will click on your link, but if you can generate a million new "views" per month, then that's 1000 clicks per month. So the key thing for a spammer is to be able to automate as much as possible. One key element in preventing spam, therefore, is to try to make it so that a human has to be in the loop somewhere.
In my case, I wrote the website; I suspect any login bot would have to be customized to my specific website to be effective. I'm sure it doesn't take much time, but you'd still have to start an automatic bot specifically for my site. Why would you do that if you saw that all of your attempts to get spam up had failed?
In your case, you're using a standard tool. Someone's already written a bot that can log into any moinmoin instance; and almost certainly someone's written tools to scan all websites for new moinmoin instances and try to create accounts. It's probably unlikely anyone has specifically thought about your site at all; they'd probably have to write special code to remove it from their automatic scans. Same thing as before -- even if their bot fails on 99% of moinmoin sites, that 1% makes it worth keeping it going.
> PS: SPAM wise the bots were just an annoyance. But every page hit runs moinmoin's Python code and it's not the fastest thing.
Golang FTW. :-D
And when you can't prove that it was the real person who paid for it, you're going to lose that one.
It is a nightmare.
I ask because we are looking at extending our platform to include enterprise file sharing and rendering for niche file types, bit don't want to expose ourselves to the extra work you've described.
Obviously there have been some issues and obviously they have processes in place to minimize and respond to abuse of their platforms, but overall both platforms rely on "trust by default" which is interesting and leads me to think that your comment might lean too far towards pessimism.
A stolen fruit is gone from the supermarket and cannot be sold to a paying customer.
A downloaded movie remains the same.
There is only a loss, if the person who pirated it, would have bought it otherwise.
So no, it is not the same.
Let's say, most people would never dare to make racist insults to anyone on the streets but many do so copiously online.
But, holy cow, every time I go on there, it is a hot mess of angry people screaming at each other. I have to assume those people just don't go outside at all?
IMO it is NextDoor's hyperlocality that makes it such a hot, angry mess. On Facebook, your cousin that's 5 states away's friend that's straight up posting swastikas unironically? The chances of meeting them, ever, is fairly close to zero if you don't want to, so it's far easier to walk away. But on NextDoor, it's because that the people are in some cases, literally next door to you, that the conversation is almost immediately emotionally threatening. It turns out your neighbors aren't like you, but not in a good way, and it challenges your sense of belonging to the local community. It turns out George down the street who's a sweet old man who's lived there for 20 years and has a pretty dog, is a neo-Nazi. Not something that would come up while exchanging pleasantries about the mail being late today, but thanks to the Internet, his eccentricities are on full display. Good fences make good neighbors.
It's easy to dismiss Nextdoor as a hot angry mess, but there are very sweet things too. My neighbor just recovered their cat, someone is collecting clothes to bring to a women's shelter, but the hot angry mess is very shouty, and impossible to tune out.
This is a excellent analogy that I don't recall having encountered before; thank you.
Most Americans for the former, maybe. "Get into a stranger's car" is literally an everyday routine in other places in the world, and has been for a long time.
In my case, image uploads aren't mandatory for users, but they are very helpful in identifying the spammers (for some reason, the spammers almost always try to upload images that get filtered, so it makes it easier to spot them). That, combined with an IP check (getipintel.net) has almost completely eliminated the spam issues on my site.
I didn't know about this service and it sounds really cool. Kind of sucks that VPN = malicious since I use Mullvad for legitimate browsing purposes. That said, I understand that VPNs are often used a medium for illicit activities.
[1] https://www.zdnet.com/article/mozilla-suspends-firefox-send-...
It reminded me a little bit of the guy that started a "buy a gift card with bitcoin" site, not realizing he was effectively creating a way to convert bitcoin into actual cash while avoiding an exchange. It was wildly successful until he realized that it was really really hard to get large quantities of bitcoin rapidly converted into cash.
However, for about 10 years, there are bots registering every day, some of them even make realistic accounts with cool and unique username, email, even description. Strangely, the email and username never match... And they make post with html links embedded. They even know to actually select the 'HTML' content type, which is an extra select input. Some bots even make a few innocent posts, before the link spam. But just too few to get past probation. Obviously, the spam posts never get approved, and accounts never get out of pre-moderation queue and yet they're still trying every day... Not intelligent enough to make more than 2 dumb posts.
Similarly, I have some work projects that have user registration with a human on-boarding process, where another person has to add the user to their group for them to share any data, none of which is ever public. So these bots are tirelessly registering, and staying in limbo forever. Thousands of useless accounts.
It boggles the mind how much energy is wasted, but I guess it must my profitable enough.
This process works well but it can also turn away a ton of people from ever joining.
I know there's this one product that only offers support through public forums and new accounts need to be reviewed by a moderator and then you're not allowed to post any threads until you've made a few replies to other threads and it's been white listed by a moderator, and then on top of that and you also can't use links until you've met some other criteria.
But in order to get proper support it requires linking to large files (videos) that aren't able to be uploaded directly through the forums.
It becomes such a pain in the ass to open a support request. It could easily take over a week just to post your question and the worst part about it is the forum software orders posts by date, and your post will be buried on the 8th page before it's visible because it takes your original pre-moderated post date as when it was created.
1: You're working to train for free to train $megacorp's image recognition or OCR algorithms.
2: Even after [what seems like] eleventy billion Captchas and eleventy billion people pointing it out to them, they STILL don't tell you how to do the damned things properly. ie. If you ask me to click on any image with traffic lights, am I supposed to only click the lights?.. or the posts as well?.. what about a square that only has a wee part of one in? Does that count, or not?
3: As above with text based Captchas. Is it case sensitive? Do I have to include spaces between the words? What about the punctuation? Or the fragment of the previous or following word, peeping in at the side?
4: The Captchas never make allowances for the end user's language settings, so I'll often get American terms used, where the thing I'm meant to be identifying isn't called the same in real English. So I'm not 100% sure what I'm looking for.
5: If you have an ad-blocker, you'll usually be asked to solve about three Captchas in a row, thus cubing the annoyance.
Thank god for 'Buster' is all I'm going to say!
[I'm not going to link to it because the less people know about it the longer it'll keep working].
One of the websites I host is for a retired elderly middle-aged academic who pens articles about his former field and invites comment.
Almost every week I'll get an email from him along the lines of "Is this worth following up?" and attached will be either a comment reading something like "I love your article. So much good info and useful to me.." or an email from some g3gergergew@gmail.com address saying "We love your site. It has much good infos and is useful but we are notice your CEO could being better..."
No matter how many times I tell him to look for the telltale signs...
* Random gibberish email from address
* Half a dozen links in the comment
* Broken English
* Generic text which could apply to ANY article or website in the bloody world
...he'll still forward them to me and ask if they're worth getting in contact with. Every time, I'm gobsmacked how he can fall for such inept spamming.
[But a discussion here on the ridiculously obvious scam that people fall for all the time is probably too off-topic --even for HN!]
PS: Upvote for "eptitude". Even if it's not a real word, it ought to be. [And I'm not going to spoil things by looking it up]
Intentional or not, they have figured out a kind of clever selection bias.
If the client doesn't stop, at least the worker is compensated highly for their trouble, so it's a win/win.
That seems like a decent goal.
This is an old person who would otherwise fall victim to a scammer, and all it takes a little bit of someones time every now and then. Suggesting we charge them $100 every time is ridiculous.
Especially young professionals struggle to understand why any of this is important, and are left wondering why they are not getting a response.
Many - most - native English speaking people don't really know 'bad' English. People I've known my entire life who've grown up in the US and went to university still have trouble writing more than a few coherent sentences. Folks are bad at writing, and I think tend to not look critically at bad writing.
It's not surprising because spammers write scripts for all sorts of platforms (WordPress, vBulletin, etc.). There's probably no custom code written to attack your site.
They detect your platform and use a script from their pre-existing library to post their spam.
For realistic usernames, they can just reuse names they've seen on other platforms.
Goes the same for emails. Any existing email list could be a source for genuine sounding names. Just throw a couple of numbers at the end of the name and you've got a unique name that a human already came up with.
I thought it was kind of funny how easy it was to defeat the spammers, but on the other hand, there are so many easy targets that it's not worth their effort to try to overcome something so elementary just for our site. Conversely, there are much higher value targets where captcha is very valuable and worth the effort to implement because it is worth it for the abusers to try to defeat them.
Seems to me there are hundreds of lifestyle businesses just waiting to happen by following this formula. So many good ideas out there could be made so much better by reducing them to their essentials, but making them elegant and "crazy fast".
Want to run a little webhosting business focused on small businesses? Tried and true. Want to offer custom software development for some little vertical? Easy enough. Want to just make a really nice and clean note app? That works too.
When your goal of success shifts from "mind boggingly rich" to "sufficient for the lifestyle I want", it's much easier to be successful.
That having been said, there are a lot of products out there that made their product intending it to be free, and then when they hit 1m users they started thinking "hmmm, if I could get a dollar out of every user, I could buy a house". They try to stuff a monetization model in sideways and damage their product in the process. Taking a moderately successful product that's crippled by attempting to shoehorn in monetization and redesigning it to have reasonable monetization from the beginning might be a better strategy.
That's exactly the point of this approach. Don't try to solve everybody's use cases like Word. Target one specific group and make the product faster and easier to use by removing all unused features.
But there are a lot of opportunities elsewhere to make products that are faster, simpler, cheaper, and more useful than the current industry standards.
I hate it when I have to use Word for something.
That sums up so perfectly what I did with a B2B SaaS product of mine. I never found words so perfectly as this quote does to describe what I was aiming to do.
I would love to take a look at your B2B project for inspiration if that's possible :)
When I've tried to cold-sell customers in the past, they often want that one feature (a random integration, feature matching from the service I'm replacing, etc.) and aren't willing to make the switch until I add it.
Maybe this wouldn't be a problem if I was marketing instead of selling?
Edit: Snap as in ditto - I would also like to know this, not Snap, the OP somehow founded SnapChat.
Wow. Traded all useablity for endless features.
Our setup barely works. And creating a story is such a massive pain. Sometimes things just won’t work.
So I refuse to use it for stories beyond one mega story.
What makes this worse is when 1/3 of the teams in your org decide that they want a similar feature, but each comes up with their own custom-named tag ("owner", "lead", "point person", "point of contact", "jefe").
And then you want to run some jql queries against those tickets, and you have to use disgusting query generators to de-dupe the tag monstrosity.
They've done an excellent job of giving you just enough features to shoot yourself in the foot with.
Maybe I should sell notepad as a fully extensible work prioritisation tracking platform that integrates with all email providers and now supports slack...
You're a UX designer. a User Experience designer. User Experience is your profession. And you're telling me: JIRA, the application with massively nested hierarchical layouts and a 2 second response time for every user interaction is easy to use?
Jira has to be the most complicated monolithic issue tracker available. I should think even Jira's sales team would struggle to call it easy to use, particularly when compared to more simple competitors like Trello.
I certainly have not made UX my specialty so I'd like to know what markers I'm missing out on here. What kind of things is Jira putting forward that mitigates its crazy abstractions, its hierarchical layouts, combined with its slow response times? I thought those were the kinds of things that signified really bad UX.
Personally. I'm forced to use Jira as part of my job and have suffered terribly at its hands. If someone is finding it a breeze then it would be a great benefit to me to understand.
That, combined with generally treating their content creators like they are completely disposable. I hope someday someone disrupts YouTube. It seems to me, besides the "network effect", the main difficulty here is unfortunately the cost of bandwidth. I could host a reddit clone from my home machine or some cheap VPS if I wanted and scale up to several thousand users, but video content at 5 megabits per second... How do you get bandwidth cheap enough to host that? Are there hosting providers that will just serve files over HTTP for super cheap?
No it didn't : https://en.wikipedia.org/wiki/ISO_8601
Specifically 4.3.11 defines the start of a decade as any year where [year]%10==0.
Also there is no day 0 of the month, no month zero of the year. It's how calendars work.
Content is a public good in the economic sense: zero marginal cost, high fixed costs, nonrivalous, and poorly excludable.
Advertising is a rent (to advertisers) and an imposition (to its audience). There's an active aversion to it, the content is very often deceptive, manipulating, and against the recipients' true interests. At the same time, as more attractive audiences are driven away (and they will be) those who remain are subject to ever more, ever lower quality ads for ever more manipulative or harmful products and services.
Bandwidth and availability must meet or address peak demands which are infrequent though often predictable. Users' decision criteria, as with highway traffic, externalise most costs, whilst benefit is privatised, incentivising overuse. Provisioning must be based not on some average utilisation, but on probability of availability and service quality, with additional nines costing orders of magnitude more to provide. That risk environment may itself be quite variable.
The consequence is that demand, revenue, and cost components follow vastly dissimmilar dynamics, making the business exceedingly difficult on market terms, and incentivising numerous pathological behaviours.
I think you have very quickly found the reason for YouTube boiling the frog. Bandwidth, encoding, and storage all have costs associated with them. They are not simply free and so YouTube has to pay for these costs somehow.
There has been a mindset shift in the Internet at some point where originally people paid for premium services, and then it swung to everything for FREE. But nothing is actually free, what happened is a tradeoff in who pays from the end consumer/producer, you, to advertising companies.
So it's now up to all of us as users of the Internet whether we are happy with that deal. I personally am not, as allowing someone else to shape my thoughts in exchange for free services is not something I believe is beneficial to humanity. I'm doing my part, but it's up to each and everyone else to make their choice and do their's. Or don't if you are satisfied with the current deal.
As for paying for content, I think there's an opportunity to combine ideas from YouTube with Patreon. Tipping or supporting specific content creators you like. You move away from an ad-driven model towards a model where you get some basic content for free, but you can pay extra for more content.
There is likely some new model that doesn't require fully ad-based or fully premium paid. We work with customers every day who are iterating on a variety of business models. I think in the end we will see a disaggregation of YouTube just like we have seen in many other once complex and costly spaces in tech.
Except it doesn't come from Google, does it? The amount of monetisation is entirely in the creator's discretion.
Disclaimer: I do work for Google, but that's public information.
Anything else requires a constant uphill battle of content filtering and deletion. You could call it censorship, but it's a necessary reality.
Depends on how you structure it. If they're paying to list, then it makes things a lot easier.
Charging nominal fees are going to keep out the most casual users but I'm not sure it's a bad idea.
Porn and spam always take over eventually, because they're consumed by, ahem, highly motivated individuals.
For example, my project is a PWA which (I assume) makes it harder for spammers to use because there are fewer direct links that could be used for spam.
What about using verified emails? Google's captcha (or similar)?
Have actual (paid) humans review the content that users upload.
Google did that with YouTube Kids, and look what horror that spawned.
Commercial spammers understand shadow banning and check for it.
The best option for commercial spammers is to make it expensive to post on your platform. That means shutting them down right away, so the effort to sign up doesn't get the reward of having an account available for days.
Mind you all roads built get used for crime. There's a point at which mitigations for advise of their services become unreasonable to expect of a company.
"I made a great fast, free porn hosting site. But then these nerds showed up and started uploading their Git repos onto it."
i.e. A network between, Messenger<->Viber<->Telegram<->Line App.
By design no media sharing was allowed(to prevent pornography) and the user profile images were received from the platform itself. But we soon faced unique challenge of people from certain countries using their children picture as profile picture(often just the children), there were people with group photo as profile image and then there were people who using explicit images as profile pictures.
So we integrated Amazon Rekognition to identify children, group and explicit images. Those using explicit images were banned immediately and those with children/group photo image(Face detection not Facial recognition) were asked to change their profile image(Their profile was not shown to anyone until they changed their profile picture with just them). We were processing >200,000 profile images per month as people change their profile images often.
But, as we very well know Amazon Rekognition or for the fact any such ML solution is not 100% accurate, we faced issues with people with darker skin color(Amazon told me that they were working to fix the issue; exactly why this type of half baked tech shouldn't be used for things which can cause harm) and so we had to reduce the confidence levels to such an extent that anything resembling a child would be flagged by the system(False positives are better in this case than false negatives).
Why would any company focusing on privacy partner with Facebook or Google? I would guess that some ardent supporters of such a product/company would be put off by such a partnership, no?
> Chat-App-Network
Messenger had >1Billion users and so its users were 98% of our user base. We enabled them to communicate with users of other chat apps and vice versa. We didn't even use the 'Name' of the users.
I applied for their bootstrap phase under their FbStart program but they directly selected it for the Acceleration phase.
As for why I applied for Facebook if I care about Privacy?
I was a disabled, solopreneur from a village in India, without any kind of network strength competing with Valley behemoths and any kind of help is not just a force multiplier but life or death(But my product was selected meritoriously by FB). Facebook privacy issues(Cambridge Analytica) started only several months after I launched the product, so the image of Facebook was not like what's today. But it did bother me and just after a year of running the platform successfully, I had to close my startup due to my health issues[1]. I did not sell my platform to safe guard the privacy of the users.
[1]https://abishekmuthian.com/i-was-told-i-would-become-quadrip...
Thank you for your kind words. I can visualise what you had to go through. I share mutual respect and wishes for your recovery or should I say 'Management of our conditions'.
>I still experience various neurological issues for which they can't trace
The main issue I had (Tingling on the Face) has been successfully resolved after the surgery. Any other discomfort I had has been largely due to anxiety and post-traumatic stress from surgery, loosing my hard earned startup etc.
So targeted efforts in bringing down the anxiety helps me a lot[Staying in present, taking lesser sensory inputs(had dozens of phone calls earlier, now its zero phone calls, only email].
Side note: Does wearing hear rate monitor on the wrist, like Fitbit/Apple watch hurt you after 15-30 mins?[1].
[1]https://abishekmuthian.com/my-experience-with-fitbit-charge-...
I do not wear fitness trackers. That's a very strange issue. I see you outlined wearing it too tight or EM sensitivity, but there's a couple of other things: heat and sweat. Both of these things are known to play unkindly with nerve damage. For example, I get pins and needles when I sit in the sun. I get a burning nerve pain when I sweat from exercise.
The article I attached has highest number of visitors since I published, several people have also asked me the reason for it. Unfortunately, it often gets dismissed as wearing it too tight and I have definitely proven that it's not the case with me.
I'm planning a blind test with an elaborate setup to conclusively prove that heart rate monitor is hurting me and others; perhaps then we could proceed to find why.
We soon discovered a similar problem to OPs - bot accounts (mostly @qq.com addresses) were registering by the hundreds per day to create wishlists and then send those wishlists to other @qq.com addresses. They were setting the titles to arbitrary code blocks.
I found it fascinating, if terribly inefficient. Some colleagues and I were speculating on the purpose, perhaps someone experimenting some kind of laundered botnet control path?
We tried all kinds of measures to prevent it but ultimately we blocked all @qq.com accounts and eventually disabled the wishlist feature altogether as it had such little real usage.
"Click here to see a translation http://whatever.r/?username"
So when we sent out the email to verify the signup the receiver saw some English text they couldn't read and the above instructions in Russian telling them to click on the link.
This was a Magento site so I assume it was a standard bot.
- One time, they used our platform to publish links to their streaming websites for the quarter finals of the 2018 Champions League. Suddenly we ended up being first result on Google for "arsenal v barcelona". It was fifteen minutes before the game so you can imagine that we got a lot of traffic. On the one hand it was kind of flattering that the SEO ranking of our domain was so strong. On the other hand, it wasn't great nor beneficial for the platform to be abused like that. As a counter-measure, we decided to block indexing of project pages for 24 hours when they're first made public. The spammers never came back.
- Another time, we got an email from AWS that our SES bounce rate was 15%, and it was rising fast. Being blocked from sending emails by AWS would have been a disaster. It turned out that our invitation system was abused. A creator of a project can invite an external person by email. That person receives an email saying "John Doe invites you to collaborate on 'A nice project about the 2018 Champions League'" with a link to the project. Replace "A nice project about the 2018 Champions League" by a Chinese ad and you've got yourself spammers who are sending thousands of emails a second to a random collection of email addresses. Naturally a lot of these bounce, which caused AWS to warn us. So we had to start verifying the MX validity of invited email addresses and throttle the system to a maximum of 100 emails in a window of 24 hours.
- We still get a lot of spammers publishing obvious spammy projects. One thing that has helped is the Clearbit Risk API. You send them an email address and it comes back with an assessment of how spammy the address is. We use it for certain domains (protonmail.com, yandex.com,...) and it frequently flags someone as a spammer right after signup. They can still use the platform but can't make stuff public, completely defeating the purpose of them being spammy.
I'm sure they'll keep finding creative ways to get around the limitations we put in. The toughest is to find a way to counter them without it hampering the experience for all the other users.
That domain name must have a negative value now?
If you're one of those loud folks who dislike instagram/facebook (like me) then this is a nice way to ensure your content and data does not end up on the platform.
Of course, they're only enforcing it themselves, so it's unlikely to be permanent. :(
> In 2018, Instagram banned the site due to "spam," although it was lifted and Instagram issued an apology.
> Rumors circulate that Instagram will issue another ban.
It's just how human nature works. 90% of people are great, but 10% can do a lot of damage.
We had outbound referral links, basically to monitor the number of times that a visitor clicks out of the website. The URL pattern was something like: example.com/out.php?url=outbound.com
The out.php script would just simply (naively) redirect the user to the url specified. We never validated if the outbound link was to an authorized reference.
The result is the same. Eventually spammers figured out the above link and then just started posting their spam using the site's redirect script to any number of social media sites, embedded in email, etc.
What's interesting too is that we would see multiple redirects embedded into a single request. e.g. out.php?url=another.link.service/link=spammer.com
Obviously in hindsight this was stupid, but when we built it (some 10 years ago), the idea seemed pretty sound if not maybe a little naive. The solution would have been to only allow redirect links to authorized outbound sites, which works when those links are relatively static and not open ended.
1) register domain
2) don't renew
Not completely unrelated either, since it seems like OP's domain is now bust too...
What is also interesting is that probably half of what FB or <insert nonexistant competitor> does is moderation of sorts. This is why FB is becoming a commerce/community page website and why they need Instagram for social media.
Hopefully I won't do anything dumb like the naïve contact form I'd made that turned into a spam vector because I was putting unsanitized user input into mail headers.
But maybe that was a game of a few years ago, and now spam bots are mostly just trying to push porn on social media.
The other thing is signup-spam. I'd say a full 50% of my signups are spam, and I do my best to remove those user accounts. What surprised me was that the spammers seem to be human, i.e. using a Gmail address (which requires verification), solving Captcha, entering form fields, clicking the email-verify link, etc. Just crazy. Again, the countermeasures are not something I would want to publicize...
email: https://mxtoolbox.com/blacklists.aspx
DNS blacklists: https://www.ultratools.com/tools/spamDBLookup
There is people who spend their entire life trying to build a working business, and those who just walk away from what would possibly be a pretty lucrative business because of lack of time.
Humans, strange and exciting beings, really.
You could allow them to select one of their accounts to source information and a photo for their combined profile. At that point you're not storing anything besides links to social media profile pages.
In effect you get to piggy-back on those sites' abuse mitigation strategies (though of course you're stuck with the lowest common denominator). Your biggest decision at that point is which social media platforms to allow onto your service.
The ToS also includes this:
> You represent and warrant that: […] the Content will not or could not be reasonably considered to be obscene, inappropriate, defamatory, disparaging, indecent, seditious, offensive, pornographic, threatening, abusive, liable to incite racial hatred, discriminatory, blasphemous, in breach of confidence or in breach of privacy; […]
For which they obviously need to look at the data.
Okay, so I have something related to this.
two paragraph back-story was cut...
So I built SASRip[1], an Open Source website with an API that allows you to download audio or video from any web-page that is supported by youtube-dl (I use youtube-dl and ffmpeg to do muxing/transcoding). I also built a browser addon for it called Media Reaper[2] (chromium version available on SASRip's website[3]).
Now, I wanted to build a no BS, no tracking website, so all I have is internal logs, these logs keep incoming requests like so: Time, URL, ID string, success/fail, the ID is just to tell where the request came from, web site, API call or the browser addon, I keep no IP data or anything like that, and boy do people download the nasty stuff, there is all kinds of nasty stuff, taboo stuff, feet stuff and stuff I didn't even know existed.
Now I live in constant fear that someday, looking trough those logs I will find CP and I am not sure what I should do, I know implementing tracking methods goes against both my morals and the philosophy of the service, but at the same time I am not sure if I can keep on going, knowing I could do something about it but I am not.
Ultimately I think it is very likely that I will shut the service down if I find CP on it, with no way of tracking the person down, perhaps leave a message with why I shut down.
----------------------------------------
P.S. I make no money on this service, it's purely donation based.
P.P.S I know I can do muxing and transcoding via some really cool JS libraries, but I wanted to sharpen some of other skills with this project.
----------------------------------------
[1] - https://sasrip.cf/
[2] - https://addons.mozilla.org/en-US/firefox/addon/media-reaper/
None of those require extra user interaction and can be quite effective.
Do you have any suggestions for further reading about approaches to this? I do plan to do manual verification initially, but if the project is succesful, I won't be able to keep up- a good problem to have, and I don't want to optimize prematurely- but I also don't want the site to become overcome with spam before I have any clue how to handle it :)
So for instance, ALWAYS do HTML sanitization via whitelist; don't let anyone put any javascript or any weird CSS or HTML into your posted content that you don't allow there.
If you allow links, make sure they all have the "nofollow" tag, so that you can't be used for link farming. (Both because it helps spammers, and because if Google detects that your site is being used for link farming, your own sites bomb.)
Other tricks spammers use: Using a link for the username, since that is sometimes emailed or displayed even when moderated content is not.
The site I run is a special-purpose conference website for a relatively small community (usually 60-100 attendees), so manually moderating all content until a user is verified works pretty well. The first year we had a handful of spammers, but their content was all deleted by me before it was seen by anyone. (With the exception of the link for a username. Missed that trick.) The next year we didn't have any spam accounts at all.
No idea how well that will scale for your use case.
Thanks.
I wonder what optimizations were done so you can only pay for 1M+ hits?
That's a really naive view of how hits are distributed over time.
Of course, hits are not homogeneously distributed. But an average of less than one hit a second leaves more than enough leeway for any reasonable clustering you may come with. Any small computer can handle thousands of times more than that.
There is some questions about bandwidth costs, that can vary wildly.
The whole app lives in Firebase. Once a user hits a profile page, a cloud function runs, get the data from the DB, constructs the page, and caches it at the CDN for 24h.
If a user add/delete something on their profile, the cache is removed. Otherwise, the content is served from cache only with no need for a lot of computation.
With this, you can still get your way through Firebase's generous free tier https://firebase.google.com/pricing
So it's not covering 1M+ invocations per month on the free tier.
1. Use email verification / captchas
2. Block mailinator/throwaway account addresses
3. Use something like Cloudflare/Akamai to protect against egregious bad actors
4. Add some other level of social validation like a major OAuth provider
5. Simple pattern matching - as the author noted, these are rarely sophisticated, and once you put even a small barrier up, they usually trickle down. You can even go further and just limit the features your product has specifically to hurt their use case, or you can actively shadow ban the users yourself, so they don't know their links aren't working.
There is meta-moderation on top of this which has users check the moderation of other users. I assume people who repeatedly abuse their privilege are given sub-naive priority for mod point assignment.
This scales with traffic and requires the minimum of admin oversight.
However of course, the larger a page gets the more dedicated the bots get.
But for the plain porn/russian spam it helped.
I've also worked at some start-ups that don't put this into their contract (or remove it when asked) so as not to scare the talent away.
I think you're right for the larger corporates though. It's often not enforced, but you run the risk of them claiming that it was "done on their time"-style issues later if its successful.
The other gotcha is that side-projects shouldn't be in any way competing directly with my employer. If you work in product-development that is pretty easy to isolate. Agencies on the other hand could basically make anything for a client. So that's a bit tricker.
No idea what of my contractual limitations are actually enforceable, though. And they've never complained about side-projects because my own projects aren't anywhere near something they would accept as a client.
E.g., in CA there are only a few whitelisted situations where your employer could own your side project, and in all other cases it belongs to you. As a matter of practice, most companies in the area allow side projects, and many "require" that you ask permission -- since they'll just rubber stamp approval onto whatever you're doing anyway, that typically works out in your favor; now you have written proof that your employer doesn't want to exercise any potential claims to your project even if it nominally seems to infringe on their business.
Of course that doesn't mean many don't try block such things…
Most companies I’ve worked for have a provision that states they can buy my projects, by compensating me for my work.
HN does the same with spam, which is the only reason we can have this conversation in the first place. I take it that you do not consider that censorship?
As for spam, it's a shame that after 20+ years of the stuff, we still don't have a good answer for it. Some subset of people is spending money on whatever spammers are advertising, or they wouldn't be doing it.
None of what you said is limited by the freedoms that you already have, including starting your own platform.
But I'm sure that when you do, you too will find some things on it that you will want to block.
> If I want to private message my friends dirty pictures, who is FB (or whatever platform) to tell me I can't?
VERY quickly becomes "unsolicited dick pics." That "links to share porn" quickly becomes "links to share child porn", etc. You're not accounting for all the genuinely abusive users.
These overarching powers to censor anything for any reason is disturbing to say the least.
Let me consider with two cases.
For one the issue of a company making the rules, e.g. Facebook imposing prudery worldwide, even in places where things like a bare breast are not considered as such (yet?!). Why should someone half a world away physically and maybe even further morally have a say in what people can or cannot post in e.g. Europe?
Another is bans/ignores/mutes by users, I would say. The tools for these are absolutely lacking. Mutes or bans are usually permanent and there is no way to do anything about it. It's like solving every problem with a sledgehammer. (If I wanted to, I could go and mute or ban anyone I like to on e.g. Twitter and they would have absolutely no means to explain themselves, appeal the decision or have any expectation that their sentence is going to be over one day.)
Maybe social media should be rather treated like public infrastructure and should have to provide services too all while leaving any moderation problems to other entities (and should just have to execute their decisions instead of both deciding and executing).
/rant
When you choose to use Instagram you buy into their censorship