Twitter Could Have Become a Protocol
blog.austingardnersmith.me
blog.austingardnersmith.me
They had an experimental project called 'annotations' where you could attach 1k of json to each tweet, like a DIY microformat.
I got onto the beta and created a prototype twitter client which you could attach mini 'apps' to tweets based on the payload type, e.g. you could tweet out a poll, or an inviation to play a game or a job advert or whatever, and you could attach your own app as a listener.
I think it could have been pretty amazing, basically A Message Bus For Everyone. If you build this yourself, you will run into the chicken and egg position, twitter could have pulled it off as they had the eyeballs and the developers.
Unfortunately they pulled the plug on it, around the time they started to close down the ecosystem.
If you run a central message bus off which people can produce their own applications, you'll find yourself having to deal with everything that ISPs and mail providers have to deal with. That means being asked to keep logs for years, dealing with DMCA takedowns, filtering the worst of the internet, etc.
And for what? As soon as anyone relying on it achieves scale they'll re-implement it themselves.
Besides, what you've described is the opposite of a protocol: an inner platform.
As for policing, I think that would probably be where the money would be made - reputation/quality of the messages being posted.
It would end up full of fake job spam, and the listeners might want quality filters. I can imagine that you might pay to validate your identity as not-a-spammer.
I'm not saying it was a slam dunk, but a lot of talk about this back in 2007/2008 and there was a palpable excitement about Twitter apps. Within a few years the media-company mindset had taken grasp, the valuations and expectations exploded, and now we have boring old Twitter which is still a damn cool thing in its own right, but which is seen as a failure because of a combination of pedestrian vision and mismanaged expectations.
That's what we're doing with "Oh By"[1]. When you speak of a "message bus for everyone", that's our goal.
You are correct, however, that there is a chicken and egg problem - Oh By Codes are interesting if everyone knows what Oh By Codes are. Otherwise, people just wonder why a weird code is written in chalk on the sidewalk ...
[1] https://0x.co
Quoting a portion of a great answer [2] from HN user niftich, that feels very appropriate here:
> "Not enough people make new running-on-TCP or running-on-UDP protocols because new protocols are hard to design, they don't work with the one application where everyone spends 70+% of their time (the web browser), and they probably get blocked on a middlebox except if you use port 80 or 443 and fake being HTTP anyway. For all but very specialized use-cases, vomiting blobs of JSON (or if you want to feel extra good, some custom binary serialization format like protobuf or Thrift or Cap'nProto or MessagePack) across HTTP endpoints is pretty okay."
Just because the traffic is encrypted, it doesn't mean the connection type isn't identifiable.
My very first job (back in 2000) actually did exactly this - we tunneled arbitrary TCP/IP traffic over HTTPS. The startup died over various management & investor relations fuck-ups, but the product was working - you could do things like make a SourceSafe connection across a firewall without exposing the server to the Internet. There was some concern about whether firewalls would kill the connection if it was open for too long, but it turned out that none of them actually did this (which unfortunately we didn't figure out until we were almost out of money and had wasted several months handling this case; see again re: management fuck-ups).
If one is using HTTP to transport your protocol (and the aforementioned websockets point isn't an issue) then your protocol should work over port 80 as well as 443. Barring any corporate / nationwide web filtering; but as others have already pointed out, that could also be an issue with 443 even with advantage of TLS.
By the way, was it Socks2HTTP you worked on? I remember that project back around the time you described and was quite fascinated by it.
HTTPS is built on top of TLS, not the other way around. You can't (passively) tell if a 443 TLS connection is HTTPS or a proprietary stream. You can take a guess based on statistics (which is what the Chinese firewall does to detect tunneling) but that's about it.
Warning: nonsensical brain dump follows:
However to address your point, you might be able to use SNI (which is sent from the client before the TLS connection is encrypted) to make some assumptions about the content. Granted this would be more in the realm of web filtering where you'd blacklist suspect domains or - in extreme cases - banned terms within hostnames. I wouldn't be surprised if SNI is one of the "statistics" the Chinese firewall uses (I'm not familiar with the implementation details of the Great Firewall of China")
1. Almost no way. The great firewall of china has traffic models of what payload sizes should look like for upload and download traffic of standard HTTP over TLS. If you start to tunnel TCP/IP over it, the pattern changes enough (small payloads for TCP ACKs, etc) that they will inject a TCP RST into the stream to screw up the connection. It's impressive and super frustrating.
While you're right that you could just use TLS without HTTP (as a great many services already do), the comment I was replying to was talking about running over HTTPS (he specifically stated HTTPS), ie TLS + HTTP. Which is effectively just using HTTP as your transport.
Like everything in IT, there's multiple ways one can approach this problem and it's probably fair to say that using TLS without wrapping your data inside a HTTP body makes greater sense if you're writing your own protocol from scratch. But on this occasion the post I was replying to - and many of the posts that preceded it - did make frequent references to HTTP and HTTPS.
> The great firewall of china has traffic models of what payload sizes should look like for upload and download traffic of standard HTTP over TLS. If you start to tunnel TCP/IP over it, the pattern changes enough (small payloads for TCP ACKs, etc) that they will inject a TCP RST into the stream to screw up the connection. It's impressive and super frustrating.
Ouch. Impressive though. Could one change the payload sizes of the TLS connection to make it look more like HTTP traffic? Shouldn't be that hard to do as most of the time you'd probably just need to add junk to the end of the server replies (assuming you have a client / server relationship with your TLS protocol). You'd probably need to make the TLS connection RESTful as well - to further mimic HTTPS and the limited connection times. Though before long you've just reinvented HTTP....
By the way, how well does websockets work over TLS? Do they throw up false positives on the Chinese firewall?
That's a weird statement to make given the point of firewalls is to limit ones access to a particular resource. If you're the systems / network administrator then the extra work is part of the job securing your infrastructure. If you're not an administrator then you're bypassing the security measures put in place by your administrators - which may well be in breach of your employment contract (as they often outline IT policies). Worse yet, if you're not even hired by whatever company nor individual who owns that infrastructure, then what you are doing is illegal.
If they agree on letting employees use a service they'll let the protocol go through the MitM. Not only that, they'll let it go through the firewall on its native port without the need of tunneling it into https.
But if a company blocks a service, employees should not circumvent the block. That would be risky.
Curiously, games using UDP still work fine. Stop justifying bad protocols with the proxy straw man.
It's a possibility, especially since Facebook became that place instead. Ultimately, Twitter didn't because they wanted to focus on driving more traffic through their first-party app (presumably as a captive client on which they can eventually display ads), and because they focused on cultivating a community (like Medium, Tumblr, LiveJournal) rather than a resource ecosystem.
[1] https://dev.twitter.com/rest/public [2] https://developers.facebook.com/docs/graph-api/reference [3] https://developers.facebook.com/docs/sharing/opengraph/using... [4] https://developers.facebook.com/docs/sharing/opengraph/objec...
The problem I see is that there doesn't seem to be any way to make money from this, for Twitter.
And really if everyone started using it this way, the privacy concerns would be even greater than the concerns people have about Facebook.
Instead the convoluted strategy by the higher-ups destroyed the entire thing. They thought they could make money with ads and "promoted content" but they somehow managed to fumble the execution of that one as well.
There is absolutely no "moat" in Twitter being a protocol.
But it's very hard to time this well. Even Facebook and their apps ecosystem didn't create enough of a moat.
They were afraid to grow beyond tweets. In life you need to grow or die. Ten years from now, they'll be fondly remembered as the AOL Instant Messenger of the 2010s.
Says who? Plenty of businesses reach a certain size and maintain healthy levels of profitability over the long term.
"Grow or die" only holds true if you receive venture capital.
If someone knocked you on the head in 2011 and you woke up today and logged into Twitter, you'd find that little has changed.
All of these technopolies will eventually be superseded by protocols. It just doesn't make sense to continue to rely on monopoly companies to provide core services. That's not to say there can't be variations or companies that build on top of core protocols etc. Just more than one, and not for the most common aspects.
However become a protocol would not have helped Twitter get there.
I think the better path is to build and then protect a captive audience. Instead Twitter saw its audience based erode away not once, but multiple times with Facebook News Feed (Twitter for news) then Instagram Video (Vine).
Twitter once had a massive captive audience and unique data. Now their competitive advantages have all disappeared. Opening up more won't help them gain an audience.
Twitter is not an aggregator of data to be passed along. Twitter is a massive un-walled global community. Though people may want access to the data as a bi product of that. The future is in cultivating community, not protocols.
I think legacy protocols such as email, IRC, SOAP...have proven time after time that there is no need for an information specific protocol.
Twitter could be more open, but I think they realized that would be giving away the keys to the castle.
A company with $1B in venture funding isn't allowed to aspire to being a $500M business. It must grow as big as the sun, or die trying.
Twitter lost big in that regard imo
Then they introduced rate-limiting ( 60 per hour, when the station generated a Tweet every 48 seconds ), application validation, phone-number verification for accounts... I gave up after that and just resorted to RSS.
Apart from the sarcasm, I'm serious on the web part, which is terrible: all around the world people are "porting" protocols to JSON+HTTP ( example: JMAP[1] ). IRC is awesome[2], but it's not GET & POST and the new kids on the block run away if it's a real protocol instead of a HTTP hack.
[1]: http://jmap.io/
[2]: https://aaronparecki.com/2015/08/29/8/why-i-live-in-irc
You could run any application on any port: IRC on port 80 with SSL would still run just fine.
You can run whatever you want however you want, but that doesn't change the reality of the enterprise environment.
Do you want your boss to see your reddit history?
But as much as I hate to admit it, Twitter's value is Protocol + Moderation. I'm not savvy with their operations but having worked with other cloud platforms, I know that any platform that has even a modicum of visibility is immediateley abused in many ways that are hard to forsee. Malicious attacks against the platform itsef are also an issue.
Email is a good example of open and federated platform, which unfortunately carries more abuse & noise than actual signal.
I still think building a protocol that incorporates a form of moderation is possible, but I'm not sure how to solve this.
Really? That is not my experience at all. There was a span of years where spam got bad, early-2000's IIRC, but my email is now mostly signal and has been since since at least 2010. I think this is partially due to better spam filters, e.g., gmail was better than average when it went public in 2007ish, but also due to the rising popularity of DKIM and SPF.
I see the abuse and noise out there on twitter and in comments sections, etc. I don't see it with email. I know people with big public personalities do get hate email and whatnot, but I thought for the average person, email was kind of a solved problem.
But to your point, maybe the same mitigation technique can be applied.
I can't seem to track down a /reason/ for this common limit. Systems were a lot smaller back in the day, but 512 is fairly easy to hit and I'd honestly expect something in the range of 1-8 KB to be the actual limit.
Though, b/c of jerks DDOSing systems, and reflection/amplification attacks, some DNS servers are requiring TCP for any packets larger than 512.
Total Length is the length of the datagram, measured in octets,
including internet header and data. This field allows the length of
a datagram to be up to 65,535 octets. Such long datagrams are
impractical for most hosts and networks. All hosts must be prepared
to accept datagrams of up to 576 octets (whether they arrive whole
or in fragments). It is recommended that hosts only send datagrams
larger than 576 octets if they have assurance that the destination
is prepared to accept the larger datagrams.
The number 576 is selected to allow a reasonable sized data block to
be transmitted in addition to the required header information. For
example, this size allows a data block of 512 octets plus 64 header
octets to fit in a datagram. The maximal internet header is 60
octets, and a typical internet header is 20 octets, allowing a
margin for headers of higher level protocols.
Note: That was published in 1981, when internetwork speeds were likely around 1 Mbps or lower.In reality you can send messages longer than this and they will be split into multiple messages over the wire - however in the US where you had (have?) to pay to receive messages it meant you would be charged for each individual message.
TL;DR; These limits may have made sense for the MVP, but as soon as most people moved to IP clients they were obsolete.
I use SMS->tweet all the time. There is also a set of commands you can use over SMS to talk to twitter. [0] Doing so makes more sense (to me) as SMS is a reliable protocol over mobile networks.
Mildly annoying was at the time there wasn't a way to see how many SMS subscribers you had. Don't know if that has changed or not, but it left us constantly wondering how many we had outside of our IRL headcount.
Essentially you can't run an advertising-less Twitter under roughly 10K servers (order of magnitude accurate -- people will want to quibble over these numbers but they won't be able to push it below 5K). I may well be forgetting something important that pushes it higher!
End result: there's no way to monetize Twitter at the scale it operates at without decentralizing it. Many of you will go "yeah, of course -- it shouldn't be a centrally controlled system in the first place." That's a fine sentiment, but many problems you can solve in a straightforward (not easy, but doable) way in a centralized system immediately become much, much harder, and now you are asking ISPs and individual users to operate these resources for free.
Take spam as one example. The main "solution" to spam we use these days is to all use a system that see enough email to power machine models that identify and filter spam out before we have to deal with it. IOW, we use centralized systems. This gets paid for with advertising!
You can fractally reinvent the system or you can just have Twitter.
Blaine even gave a really interesting presentation I remember watching on the subject, about how very challenging such a backplane is compared to a more simple human-human messaging network.
It's really the only thing he said I thought was smart, so it stuck with me.
That method works fine for many different tools. But with social apps you need to be able to federate, which means you need to have a proprietary vendor who wants to help, which isn't going to happen.
I'm hoping Matrix will provide the bridging required to be able to send notifications from Diaspora to Twitter (for example). I will always be a hack though. :/
We, hackers, think that everything is so sexy with P2P, federation and decentralization in general, but I don't think 'normal' people share that sentiment.
People love brands, they're surrounded by them, and they feel loyalty towards them. If there's no company behind something, it just won't feel right. You won't feel that push towards a thing. The reason why people use Snapchat isn't just because it's a good tool, it's also because it's cool.
I just don't see how something like GNU Social will ever become cool, if there isn't some timely and powerful brand to push it.
The thing is also that we don't have any clear incentives to use P2P-services right now, because the silos aren't posing any clear threats to us as individuals. If we see something like a major data breach, or something like a "Snowden for social networks" that change the way we relate to these behemoths, we might just see users getting ready to give them up.
[1]: https://mastodon.social [2]: https://github.com/Gargron/mastodon
Disclaimer: I'm the developer
(Don't become an software architecture astronaut. IRC is also a terrible protocol. It was still successful)
e.g. https://bugzilla.mozilla.org/show_bug.cgi?id=687798
There's also an IRCv3 section on blog.irccoud.com but I can't access it from here.
Chat history makes chat more usable, and it makes it usable at all on intermittent connections.
As far as i can tell it would add no benefit but only overhead to the protocol.
It's like asking why would http protocol allow users to send data. Receiving is enough.
I like IRC because its simple. I've build a IRC out of boredom, and a bouncer because the one i used missed a feature i wanted. Please nobody take that simplicity away :/
Please, don't do that! Besides, I'm not going to make a subscribe decision until I've read the whole thing.
https://en.wikipedia.org/wiki/OStatus https://www.w3.org/community/ostatus/
Never went anywhere though, AFAICT
http://www.welivesecurity.com/2016/08/24/first-twitter-contr...
Also tweets from my fridge.