3,789 karma · joined December 9, 2009
“Hey, look at Cloudflare, should we reduce egress costs?”
“Absolutely not, but let’s make a free tier change and hope that fewer people hassle us about that”
It’s a signal that they’re digging in on the price structure.
NIMBYs at it again.
> The highest levels of technology are not necessarily full autonomy, but situated autonomy
> All autonomous systems are joint human-machine cognitive systems
Fundamental questions: Where are the people? Which people are they? What are they doing? When are they doing it?
This might be a bit off topic, but speaking of gaps in common observability tooling: is an OLAP database a common go-to for longer-timescale analytics (as in [1])? We're using BigQuery, but on ~600GB of log/event data I start hitting memory limits even with fairly small analytical windows.
In this context I have seen other references to: Sawzall (google), Lingo (google), MapReduce/Pig/Cascading/Scalding. Are people using Spark for this sort of thing now? Perhaps a combined workflow would be ideal: filter/group/extract interesting data in Hadoop/Spark, and then load into OLAP for ad-hoc querying?
Even attacking SHA2-256/128 would be quite difficult as I understand it, even though it's the same length as MD5.
Truncated hashes also of course have the great property that they mitigate the length extension in Merkle-Damgard
What does scale is using fire to fight fire -- prescribed burns, and allowing natural fires to burn while monitoring them for safety. In parallel, lots can be done to harden home in the Wilderness-Urban-Interface.
https://www.latimes.com/california/story/2021-09-18/how-pres...
In practice, knowledge management at companies is a specialization. There are <5% of employees that go around and document/organize things for everyone else. Most employees are passively consuming information and information hierarchies built by someone else.
If you're not building tools for those power users, you're not building for creating and organizing content in your system at all.
As an example of how nuts this is, managers at my company regularly try out various search terms, create index documents, and do "internal SEO" to optimize how other employees will discover documents. This isn't a byzantine environment like public web search is, why do I have to hack around the wiki's default notion of page relevance?
Even early Google had more power user features than a typical B2B product search bar.
Boolean expressions (NOT, OR, AND), exact match strings, links-to, linked-from, in-folder/category, etc. should be mandatory for these workflows. Better if you can include search queries as live page content, as in Notion & Height.
If you take as an axiom that mainstream businesses will be forced to protect themselves against contributing to CSAM distribution, then the choice to offer encrypted cloud storage to non-technical end users REQUIRES doing the scan on a trusted computer (Apple has chosen the phone itself).
I think the arguments that this can be abused are very real, but it's worth talking about how to fix that, because I think the alternative might be sacrificing E2EE cloud storage in the mainstream (as has happened with every other mainstream company). Perhaps more thought should be put into making this process auditable by the device owner (or by a trusted 3rd party -- say the EFF).
Or perhaps the scanning could be federated -- say I don't want Apple doing that, but I might trust a privacy oriented non-profit to "certify" to Apple that my personal photo album is CSAM-free. Can that 3rd party scan be blinded, such that I send data that is representative of my images, but I've already anonymized my photos using a transformation?
Could we audit (similar to certificate transparency):
1) What data from the device is being scanned? What data is being uploaded?
2) What "hashes" are being matched against, and how are those changing over time? Can the data lineage of the NCMEC database be audited? Would that pick up malicious hashes injected into the database?
Generally, I think our privacy paradigm needs to be built in such a way that it can actually be deployed in our policy environment. More realpolitik, less ethical grandstanding.
If you can make an innocent picture collide with a CSAM picture, presumably you can also edit a CSAM picture to have a random hash not in the database?
SQLite: SQL -> SQLite C query engine -> POSIX Filesystem
Absurd-SQL: SQL -> (WASM) SQLite C query engine -> "filesystem" shim code -> IndexedDBAs an example: on a very short term basis, it is easier and cheaper to pile your trash in your basement, or to not brush your teeth. The value of those things only becomes very apparent after a day or more.
On a short term basis, fossil fuel based solutions are really great, flexible, stable, and cheap. We have to do something different despite that, because that short term basis doesn't tell the full story.
As an example: some situations call for being very detail oriented (J), where some situations call for being big picture (P). If you know you are prone to one or the other, you can either let others handle those situations, or realize that you are outside your comfort zone and be extra intentional.
Rules:
- opt in: nothing happens unless I click the “I saw something on a billboard” button
- no animation or sound, just a button/link
This blog has seen 14,000 views today, but some of the other pages are ~38-40k views
So if you assume ~990KB average page size, 30k views, and he were to post once a day, then that is $78/mo in bandwidth costs assuming 9c/GB. Maybe people tend to click around to the homepage / old articles, or refresh the page a few times?
Fish manages to pull this off without _any_ configuration burden (I use it completely stock), which is exciting because usually to benefit from shortcuts or macros or other "productivity hacks" you have to become a power user.