Connect() – a new API for creating TCP sockets from Cloudflare Workers
blog.cloudflare.com
blog.cloudflare.com
Over time it seems more and more confirmable as a hypothesis.
But looks like Discord Gateway blocks CF Workers: https://github.com/discord/discord-api-docs/issues/6145#issu...
Not being able to meaningfully use any external services that didn’t support an HTTP / fetch API was one of the biggest consistent pain points.
Arguably it was the one with the biggest negative architectural ramifications. Given how long (understandably) it has taken to move D1 forward in the ways that matter most (e.g. transaction support), this is a huge step towards production viability for a more diverse range of products.
When I left my company in April I had Cloudflare as a “glad I tried it, but not ready for production use / that was a mistake” - this week has it back on my list for evaluation on whatever I do next.
Congrats to the Cloudflare team! I admire your intuition for what customers need and your willingness to compete with yourself on stuff like this (actively support other DB providers while building D1 - respect).
Something I quite like doing is a thread-local (or async-local) context transaction, and that seems quite hard to do if not impossible with both batching and stored procedures from what I've seen.
What I really wish for is to drop in any old query builder or ORM and use it identically to how I would with SQLite. I'm not sure if that's feasible, however.
This is all fine when the application is using SQLite as a local library since any particular transaction can finish up pretty quick and unlock the database for the next writer. But D1 allows queries to be submitted to the database from Workers located around the world. Any sort of multi-step transaction driven from the client is necessarily going to lock the database for at least one network round trip, maybe more if you are doing many rounds of queries. Since D1 clients could be located anywhere in the world, you could be looking at the database being write-locked for 10s or 100s of milliseconds. And if the client Worker disappears for some reason (machine failure, network connectivity, etc.), then presumably the database has to wait some number of seconds for a timeout, remaining locked in the meantime. Yikes!
So, the initial D1 API doesn't allow remote transactions, only query batches. But we know that's not good enough.
To actually enable transactions, we need to make sure the code is running next to the database, so that write locks aren't held for long. That's complicated but we're attacking it on a few different fronts.
The new D1 storage engine announced a couple weeks ago (which has been my main project lately) is actually a new storage engine for Durable Objects itself. When it's ready, this will mean that every Durable Object has a SQLite database attached, backed by actual local files. In a DO, since the database is local, there's no problem at all with transactions and they'll be allowed immediately when this feature is launched.
But DO is a lower-level primitive that requires some extra distributed systems thinking on the part of the developer. For people who don't want to think about it, D1 needs to offer something that "just works". The good news is that the Workers architecture makes it pretty easy for us to automatically move code around, so in principle we should be able to make a Worker run close to its D1 database if it needs to perform transactions against it. (We launched a similar feature recently, Smart Placement, which will auto-detect when a Worker makes lots of round trips to a single back-end, and moves the Worker to run close to it.)
Sorry it's not all there yet, but we're working on it...
Intuitively I had a rough idea around why this was both a) a blocking issue ("why can't we just YOLO-try some version of it anyways?" came up at least a couple times on our side internally during alpha evaluation) and b) a hard issue, but I didn't know the details re: SQLite being single-writer specific (at least for the time being).
Valuable for my own knowledge, in addition to being useful re: understanding the steps involved for Cloudflare to enable this in the future.
I was already excited about Smart Placement regardless, but now doubly so knowing it is adjacent to what will enable a more generic solution for D1. We used DOs very heavily on the project I mentioned in my previous post, but as you call out, they are more complex to reason about, and in practice it limited who could work on them effectively on our team.
I always appreciate your comments in Worker / DO-related threads here on HN and have found them very insightful / helpful in learning more about what's under the hood! Thanks for taking the time to continue posting here - I know there's lots to do elsewhere.
This new raw TCP connection feature will undoubtedly be used to attack other services in similar ways.
Cloudflare has (had?) a murky history with not taking down DDoS for hire services ironically hosted behind cloudflare. But while you could argue they had an incentive to do that (sell protection), I can't think of any incentive to let Workers be abused.
Trivial based on the fact that HTTP requests coming from CloudFlare Workers has a cf-worker header. Also, any traffic coming from cloudflare-owned IP blocks clearly belongs to cloudflare and can be safely blocked.
Cloudflare reserves IP ranges just for Private Relay: https://developer.apple.com/support/prepare-your-network-for...
Well no, not if you yourself are also using Cloudflare
cf.worker.upstream_zone ne "" and not cf.worker.upstream_zone in {"aimoda.workers.dev" "ai.moda"}
Link: https://developers.cloudflare.com/fundamentals/get-started/r...
cf-worker: example.com
I'd love to design a database for this environment (if you're reading this at Cloudflare, you can hire me to work on this.) I think something that distinguishes between write and read requests, moves writes close to the leader server hosting the data being written, handles reads at the edge, and replicates the deterministic request itself would perform the best and give sequential consistency. This is the approach taken by fauna.com, and it's competitive with Spanner but without the need for GPS and atomic clocks to provide an accurate time source.
Opening a raw socket from a worker combined with a basic HTTP implementation could let you create a dynamic proxy that uses Cloudflare's worker IP range as the source address. That sounds like a fun^Winteresting way of getting around rate limits.
Related, does waiting for I/O count as "cpu time"? A proxy request might 10s of ms in total, but most of that will be waiting for packets to flow back and forth.
Yep, see: https://github.com/zizifn/edgetunnel
> Related, does waiting for I/O count as "cpu time"?
No.
> but most of that will be waiting for packets to flow back and forth.
If you don't await on the contents but simply pipe the streams in and out, you should be well within CPU-related limits.
Just did some brief testing. For me, the source IP wasn't consistent, but it was from an IP range belonging to Cloudflare. Notably, however, it wasn't from one of the IP ranges listed at https://www.cloudflare.com/ips/ unlike requests made from a Worker via `fetch()`. So, if you initiate two requests from the same worker -- one with `connect()`, and the other with `fetch()` -- then the first request uses a source IP belonging to Cloudflare but not from a range listed on their IP range page, while the second uses a source IP from a range listed on their IP range page.
I suspect the reason for the different behavior is to do with the `Cf-Worker` header that `fetch()` adds, which enables applications to differentiate requests made by a Worker from requests made by Cloudflare itself. Raw TCP sockets can't add headers, so they need to differentiate themselves another way.
What I'm not fond of is company's fixation only with big clients and leaving out a serious effort to bring on the average solo programmers/entrepeneurs (something on which companies like Stripe instead thrived on).
Cloudflare really needs to do more to target small fishes, they are tomorrow mid and big fishes trying to understand how to make small and medium problems trivial leveraging their platforms.
"You want to use XYZ? Fine, we include any count of XYZ for free. Next agreement negotiation. 200$ for each instance of XYZ."
CF is simply not the service to use for any large site.
I needed logging to track down a few issues with my website, but logging is apparently a feature for Enterprise only, and requires a recurring four-figure cost. Thus, I switched over to Cloudfront, which lacks in some security features and is insanely expensive past 1 TB, but at least provides features without having to pay a huge amount upfront.
I realize this is probably untenable considering certain compromises made in the CF infra (i.e. V8 optimizations), but one can dream. For now, Azure appears to be my prison.
Cloudflare has teased about containers in the recent past (but not quite running them on the edge, as it were): https://blog.cloudflare.com/containers-on-the-edge/
I believe, their very new Browser Rendering service is an example of one such deployment? https://blog.cloudflare.com/browser-rendering-open-beta/
The trouble with "containers on the edge" is that if we just literally put your container in 300+ locations it's going to be quite expensive.
Cloudflare Workers today actually runs your Worker in 300+ locations, and manages to be cost effective at that because it's based on isolates rather than containers.
We'll probably offer some sort of containers eventually, but it probably won't be oriented around trying to run your container in every location. Instead I'm imagining containers would come into play specifically for running batch jobs or back-end infrastructure that's OK to concentrate in fewer locations.
(I'm the tech lead for Workers.)
What does the roadmap look like around establishing some sort of multi-tier architecture within the CF product stack?
I imagine I could hack something together today by combining CF workers and another hyperscaler to run my .NET workload (TCP connection definitely helps with this!), but I think that there would still be a lot of friction with operations, networking, etc at scale. Ideally, workers and backend would be automagically latency-optimized and scaled relative to each other.
https://github.com/aimoda/cloudflare-worker-to-aws-lambda-fu...
We don't have this in any public example yet, but here's our simple trick with Workers on how to bypass needing to pay for Amazon API Gateway or CloudFront but still get routed to the nearest AWS location:
1. Add add Lambda Function URLs as records in a latency based record on Route 53. (Lambda Function URLs do not support custom domains, so you cannot use this record directly.)
2. Have the Worker do a fetch to `https://cloudflare-dns.com/dns-query` on the Route 53 CNAME to discover what lowest latency Lambda Function URL hostname is.
3. The Worker can then fetch the Lambda Function URL using the discovered hostname.
I eventually just decided for those use cases I'd rather run a dedicated server. Pricing is easier, no lock in, and it doesn't die on me unexpectedly.
This necessitates some more careful design - clients must assume the server can go away at any time and therefore work in a local-first fashion. I was interested in doing that anyway, so it didn't feel like an imposition.
I think the core of what I really like about DO's is that it's a single thread with a globally unique and accessible address.
Another annoyance is how little it gels with the rest of Pages. I'd expect that if I'm developing an application from the ground up in Pages, incorporating the various Cloudflare value-ads should come relatively naturally. Instead I need to have entirely separate projects just for these components.
I forgot to mention about pricing. We've taken the approach of estimating a "worst case" time per object and requests/second rate. E.g. we assume a worst case of 30 req/s for the lifetime of the project, based on 3 users simultaneously sending updates at a rate of 10/sec each. And maybe we assume 1-2 hours of cumulative time spent in the editor per object. Multiply that out to get a value of GB-seconds per object, multiply by the respective prices, and add to get the total "expected lifetime cost per thing".
(We're not actually using durable objects' built-in storage, so this would indeed be complicated by trying to price that in. We only wanted DOs for coordination.)
If it's hard to estimate the total lifetime of a thing, you could estimate against something like how many hours can/will each user spend in your app per month, and come up with a monthly benchmark price. We haven't done that as the objects we deal with have naturally limited lifespans, and our pricing is tied to the number of these objects created anyway.
But stale-while-revalidate does not work. And there is no request collapsing. And you still have to write code to get CORS to work in workers. And you can't cache the response from workers.
How does CF prioritize things? :)
while true; do pkill -f miniflare-dist; npx wrangler pages dev public; done;
Lest ye end up with a million rogue processes spinning down your CPU after each file saved with a syntax error causes the entire dev launcher to crash and leave its spawn everywhere...For incoming emails, Cloudflare Email Workers are $0 per email and SendGrid is also $0 per email. We only pay for the cost of the Worker with this setup.
2. Would not pay for click tracking. Might use it if it’s free.
3. We probably won’t use MailChannels again yet since we want to use AMP emails and ticket #221628 was handled pretty poorly. (We made our own MailChannels-like API for AWS SES after the frustrating experience.)
[0]: https://blog.cloudflare.com/node-js-support-cloudflare-worke...
[1]: https://blog.cloudflare.com/workers-node-js-asynclocalstorag...
[1] https://developer.mozilla.org/en-US/docs/Web/API/FetchEvent
Their explanation:
"the time value returned is not the current time. Date.now() returns the time of the last I/O. It does not advance during code execution"
https://developers.cloudflare.com/workers/learning/security-...
Pass.
The platform that VPSes kubernetes uses run on though, all the big clouds have a proprietary one.
I sorta stopped caring about grammatical errors since I realized English is but a second language to many people.
I wonder of The Editor has regretfully gone the way of the dodo in 'technical' writing..
if*