Is removing the dependence on US tech easy for the EU? No, it's tough and takes a lot of work and time. It's still a piece of cake compared to the dependence on Chinese manufacturing. They're incomparable.
Is removing the dependence on US tech easy for the EU? No, it's tough and takes a lot of work and time. It's still a piece of cake compared to the dependence on Chinese manufacturing. They're incomparable.
As long as mobile os and adjacent services like the store etc are controlled there is no true path to digital independence especially in a highly digitalized region like the EU.
One example is if EU allows the Android developer verification to pass this year in its current or even in more relaxed form, that just means EU is still open for some hard lessons in the future.
While a massive endeavor, it's absolutely doable to create the EU's own OS and store. It's not doable to create the manufacturing capacity needed to produce all the hardware that goes into smartphones at scale.
China itself ironically serves as a great example - they have their own Android store, mostly run on Chinese phones, some on non-Android OS. Yet they still haven't been able to get rid of the dependency on TSMC/ASML. They're working on it and will get there, but it's taking many years longer than the software part. And not for lack of trying. The fact that they're still tolerating iOS doesn't disprove the existence of the former ecosystem. iOS is said to have maybe 20% market share in China.
I find this highly optimistic. It will take years, maybe decades for EU to replace US clouds and tech. And if they're going to do it with LLMs, then it will take billions of euros in devs and tokens (again, all going to US tech companies).
Meanwhile, USA continues to strategically re-home TSMC to Arizona whilst simultaneously make huge investments to invigorate Intel and Micron.
Over the last decade USA and China have doubled-down on massive investments to out-compete each other while the EU seems like it's struggling to understand where to even begin.
Oh don't worry, Trump's already kneecapped both of those for a decade to come from 2025's actions alone. Y'all got time to catch up.
China, much scarier. But we all kinda let that happen over 30 years. Too late to complain now. I'd say we work together but uhh... I think we both understand (or rather, fail to understand) modern US policy these days.
But what’s funny is that Claude Code is from US company so can’t be used in a boycott scenario
As a compute engine its SQL capabilities are worse than the slowest pretend timeseries db like Elasticsearch.
As far as S3, are you trying to ingest a lot of small files or one large file? Again Redshift is optimized for bulk imports.
Clickhouse, even chdb inmemory magic has better S3 consumer than Redshift. It sucks up those Kinesis files like nothing.
Its a mess.
Not to mention none of its Column optimizations work and the data footprint of gapless timestamp columns is not basically 0 as it is in any serious OLAP but it is massive, so the way to improve performance is to Just align everything on the same timeline so its computation engine does not beed to figure out how to join stuff that is Actually time Aligned
I really can’t figure out how anyone can do seriously big computations with Redshift. Maybe people like waiting hours for their SQL to execute and think software is just that slow.
https://docs.aws.amazon.com/athena/latest/ug/ctas.html
I have a sneaking suspicion that you are trying to use Redshift as a traditional OLTP database. Are you also normalizing your table like an OLTP database instead of like an OLAP
https://fiveonefour.com/blog/OLAP-on-Tap-The-Art-of-Letting-...
And if you are using any OLAP database for OLTP, you’re doing it wrong. It’s also a simple “process” to move data back and forth between Aurora MySQL or Postgres by federating your OlTP database with Athena (handwavy because I haven’t done it) or the way I have done it is use one Select statement to export to S3 and another to export into your OLTP database.
And before you say you shouldn’t have to do this, you have always needed some process to take data from your normalized data to un normalized form for reporting and analytics.
Source: doing boring enterprise stuff including databases since 1996 and been working for 8 years with AWS services outside AWS (startups and consulting companies) and inside AWS (Professional Services no longer there)
Why are you doing this manually? There is a built in way of doing Kinesis Data Streams to Redshift
https://docs.aws.amazon.com/streams/latest/dev/using-other-s...
Also by default, while you can through Glue Catalog have S3 directly as a destination for Redshift, by default it definitely doesn’t use S3.
There is no need for Athena, Redshift ingestion is a simple query that reads from S3. I dont want to copy 10TB of data just to have it in 1 file. And yes, default storage is a bit better than S3 but for an OLAP database there seems to be no proper column compression and data footprint is too big resulting in slow reads if one is not careful.
I mentioned clickhouse, data is obviously not OLTP schemed.
I don’t have normalized data. As I mentioned, Clickhouse consumer goes through 10TB of blobs and ends up having 15GB of postprocessed data in like 5-10 minutes, slowest part is downloading from S3.
I am not willing to pay 10k+ a month for something that absolutely sucks compared to a proper OLAP db.
Redshift is just made for some very specific, bloated, throw as much software pipelines as you can, pay as much money as you can, workflows that I just don’t find valuable. Its compute engine and data repr is just laughably slow, yeah, it can be as fast as you want by throwing parallel units but it’s a complete waste of money.
I think these systems are optimized for something else, probably organizational scale, predictable low value workloads, large teams that just throw their shit at it and it works on a daily basis, and of course, it costs a lot.
My experience after renting a $1k EC2 instance and slurping all of S3 onto it in a few hours, and Redshift being unable to do the same, made me not consider these systems reliable for anything other than ritualistic performative low value work.
I need fast consumers, I need good materialized views.
I am not treating anything like OLTP databases, my opinion on OLTP is even harsher. They can’t even handle the data from S3 without insane amounts of work.
I do not even think in terms of OLTP OLAP or whatever. I am thinking in terms of what queries over what data I want to do and how to do it with the feature set available.
If necessary, I will align all postgresql tables on a timeline of discrete timestamps instead of storing things as intervals, to allow faster sequential processing.
I am saying that these systems as a whole are incapable of many things Ive tried them to do. I have managed to use other systems and did many more valuable things because they are actually capable.
It is laughable that the task of loading data from S3 into whatever schema you want is better done by tech outside of the aws universe.
I can paste this whole conversation into an LLM unprompted and I don’t really see anything I am missing.
The only part I am surely missing are nontechnical considerations, which I do not care about at all outside of business context.
I know things are nuanced and there’s companies with PBs of data doing something with Redshift, but people do random stuff with Oracle as well.
Did you research how you should structure your tables fir optimum performance for OLAP databases? Did you research the pros and cons of using a column based storage engine like Redshift to a standard row based storage engine in an traditional RDMS? Not to mention depending on your use case you might need ElssticSearch.
This if completely a you problem for not doing your research and using the worse possible tool for your use case. Seriously, reach out to an SA at AWS and they can give you some free advice, you are literally doing everything wrong.
That sounds harsh. But it’s true.
I can change sort order, same as in Redshift with its sort keys, to improve compression and compute. Redshift does not really exploit this sort-key config as much as it could.
My own assessment is that I'm extremely skilled at making any kind of DB system yield to my will and get it to its limits.
I have never used Redshift, Clickhouse or Snowflake with 1 by 1 inserts. I have mentioned S3 consumers (a library or a service, optimized to work well with autoscaling done by S3, respecting SlowDown -- something Redshift itself is incapable of respecting -- and achieving enormous download rates -- some of the consumers I've used completely saturate the 200Gbps limits of some EC2 machines at AWS). These consumers cannot be used in a 1-by-1 setting, the whole point is to have an insanely fast pipelining system with batched processing, interleaving network downloads with CPU compute, so that in the end, any kind of data repackaging and compression is negligible compared to download, so you can just predict how long the system will take to ingest by knowing what your peak download speed is, because the actual compute is fully optimized and pipelined.
Now, it might just be Redshift has bugs and I should report them, but I did not have the experience of AWS reacting quickly to any of the reports I've made.
I disagree, it's not a me problem. I am a bit surprised after all I've written that you're still implying I want OLTP, am using the wrong tool for the job. There are just some tools I would never pick, because they just don't work as advertised, Redshift is one of them. There are much better in-memory compute engines that work directly with S3, and you can create any kind of trash low-value pipelines with them, if you reach mem limits of your compute system, there are much better compute engine + storage combos than Redshift. My belief is that Redshift is purely a nontechnical choice.
Now, to steelman you, if you're saying:
* data warehouse as managed service,
* cost efficiency via guardrails,
* scale by policy, not by expertise,
* optimize for nontechnical teams,
* hide the machinery,
* use AWS-native bloated, slow or expensive glue (Glue, Athena, Kinesis, DMS),
* predictable monthly bill,
* preventing S3 abuse,
* preventing runaway parallelism,
* avoiding noisy-neighbor incidents (either by protecting me or protecting AWS infra),
* intentionally constrained to satisfy all of the above,
then yes, I agree, I am definitely using the wrong tool but as I said, if the value proposition is nontechnical, I do not really care about that.
Yes an according to my assessment I’m also very good in bed and extremely handsome.
But there is an existence proof seeing that you are running into issues yet millions of people use AWS services and know how to use the right tool for the job
I’m not defending Redshift for your use case, I’m saying you didn’t do your research and you did absolutely everything wrong. From my cursory research of Clickhouse, I probably would have chosen that too for use case
Your assessment of me is flawed. You haven't really shown any kind of low-level expertise on how actually these systems work, you've just name dropped OLTP OLAP as if that means anything at all. What is Timescale (now TigerData), OLTPOLAPBLAPBLAP? If someone tells you to use Timescale, you have to figure out how to use it and make the system yield to your will. If system sucks, it yields harder, if system is well designed, it's absolutely beautiful. For example, I would never use Timescale as well, yet you can go on their page and see unicorns using it. I have no idea why, but let them have their fun. There's successful companies using Elasticsearch for IoT telemetry, so who am I to argue I wouldn't do that as well.
There's nothing wrong with using PostgreSQL for timeseries data, you just need to know how to use it. At some point, scaling wise, it will fail, but you're deciding on tradeoffs.
So yes, my assessments have a good track record, not only of myself, but of others as well. I am extremely open to any kind of precise criticism and have been wrong bazillion times and I take part in these kinds of passionate discussions on the internet because I am aware I can absolutely be convinced of the other side. Otherwise, I would have quit a long time ago.
Massive endeavor for a lot of setups.
Not depending on Chinese manufacturing is borderline impossible even if you are starting from scratch. Not only it will be way more expensive, with potentially longer delays and lesser capacities, but just finding some company that can and wants to do the job can be a nightmare. From what I have seen, many local manufacturers in the US and Europe are really there to fulfill government contracts that requires local production.
Most hardware kickstarter-like projects rely on Chinese manufacturing as if it was obvious. It is not "find a manufacturer", it is "go to China". Projects that instead rely on local (US/Europe) manufacturing in order to make a political statement have to to though a lot of trouble, and the result is often an overpriced product that may still have some parts made in China.
A large corporation just migrating from everything hosted on VMs can take years.
And if you are responsible for an ETL implementation and working with AWS and have your files stored on S3 (every provider big and small has S3 compatible storage) and your data is hosted on Aurora Postgres, are you going to spend time creating a complicated ETL process or are you going to just schedule a cron job to run “select outfile into S3”?
And “most” of the services on AWS aren’t based on open source software and you still have to provision your resources using IAC and your architecture. No Terraform doesn’t give you “cloud agnosticism” any more than using Python when using AWS services.
Are you going to tell your developers to spend weeks writing ETL code that could literally be done in an hour using SQL extensions to AWS?
Are you going to tell them not to use any AWS native services? What are you going to do about your infrastructure as code? Are you going to tell them to set up a VM to host a simple cron job instead of just using a Lambda + Event Bridge?
And what business value does this theoretical “cloud agnosticism” bring - that never is once you get to scale.
It took Amazon years to move off of Oracle and much of its infrastructure still doesn’t run on AWS and still uses its older infrastructure (CDO? It’s been a while and I was on the AWS side)
I have yet to hear anyone who worries about cloud agnosticism even think about the complexity of migrations bring at scale, the risk of regressions, etc.
While I personally stay the hell away from lift and shifts and I come in at the “modernization” phase, it’s because I know the complexity and drudgery of it. I worked at AWS ProServe for 3.5 years and I now work as a staff consultant at a 3rd party consulting company.
This isn’t me rah rahing about AWS. I would say the same about GCP, Azure, the choice of database you use, or any other infrastructure decision.
>And what business value does this theoretical “cloud agnosticism” bring - that never is once you get to scale.
The "business value" here is not being beholden to an increasingly hostile "ally" who owns the land these servers operate on. If you aren't worried about that, then there is no point in doing any of this.
But if things do escalate to war, there's a very obvious attack vector to cripple your company with. Even if you're only 20% into the migration, that's better than 0%.
I of course don't know the scale of your company and how much they even wanted to migrate. Those are all variable in this.
How often has been replacing Chinese tech manufacturing dependency at scale done before? About 0.