HNHacker News
TopNewBestAskShowJobs

dhx

2,947 karma · joined October 19, 2011

-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA512

David Hicks (dhx)

Web (IPv6 and IPv4): https://david.hicks.id.au

E-mail (IPv6 and IPv4): david@hicks.id.au

PGP: Public Key: https://david.hicks.id.au/pgp/728F3435.asc Key ID: 728F3435 Fingerprint: 2442 14B5 2E51 CB0F CA3B 9DB8 59E0 E7B7 728F 3435

Hacker News: https://news.ycombinator.com/user?id=dhx

Last updated: 2012-03-05

-----BEGIN PGP SIGNATURE-----

iQIzBAEBCgAdFiEEJEIUtS5Ryw/KO524WeDnt3KPNDUFAlqcjdcACgkQWeDnt3KP NDUoJQ//TTN8pul74acpaVWImkdxEzqU1xOtAn/JyrZtYLjO2XnMDJSF9fTYKD7o AT19lV6kZnMo0p0lsNg5tnk5QMJ96rqPhpelAiXBYkdfz3VfVVj3X3AOdkg5XRbm 6Uy1sa8Omp6Jo21wtUOE+jtSYWwEiUFn9NPl3B1aNZSUo/LuMTJve1IidJuW705m OCenqn/Ro8eZ6kUOUByWoDBbGlobAkZpBZEyzHCpczZ1nGj50mn4cssXxDG4St62 S+o9009ZMgjFtfIDTwI7gwaNCgn6ldBbwYwfMaPL2cs2Kxl05iqbM/KcsU/CLZE5 Dmh/0aPaYNJdg8PDjpItupxryNwwROB/lCigGsp+msiXtOJTENdCfojFF6ffiFTl Lxqfmbe+bn3JcqTOXzExPD0NXstx7sdetQZSwVnuCQGH4JjcUpyzKmrEw4ATvpXb VXWf00jXAMWhGHSQJLzecTXw8/iWWh0swIpLrokq3ISrMBwK2oAavZzPh4TpH5ao Y+VL0LkqAaiEpLyBAJDUFFdrJbDv9P82fZ8NFEiK3plgm5xfv62UOzFVIG+ezdDY PvfodliBtu/mLMc9YIjo5H59NtnmT4fXb8/LW5rNP7xSxWjMbCGiG3/pVSUx+FBN v74Ykn2ioBPNM3ApFC3kBrlS1K8jIcNgCfMQcH0oc7y6+e84NwY= =sBF3 -----END PGP SIGNATURE-----

submissionscomments
dhx··on Fuck Android Developer Verification Program
See for example Rosado et al. v. Blanche et al., 1:26-cv-01532[1] which is an active case addressing whether the US federal government suppressed the first amendment rights of the "Eyes Up" application developer by pressuring Apple to remove this application for crowd-sourced recording/reporting of ICE incidents. An injunction has been granted with the following finding:[2]

"the Court finds that Plaintiffs are likely to succeed on the merits of their claim that Defendants violated their First Amendment rights through coercion of Facebook and Apple."

My reference to a press freedom index is due to a lack of "countries by freedom of software development" index. If such index existed, the methodology may have some commonality with a press freedom index.

Placing barriers in front of publishing of books, online content, software, etc all leads to some suppression of speech. The question is whether the benefits of the barriers placed outweigh the risks/negatives. In this case, why should anyone care what the passport number or residential address is of a student interested in a career in software development who is writing their first Android game?

[1] https://www.fire.org/cases/rosado-et-al-v-blanche-et-al

[2] https://www.fire.org/sites/default/files/2026/04/Memorandum%...

dhx··on Fuck Android Developer Verification Program
Google employees can pat themselves on the back for a job well done putting the same hurdles in place for software development that apply in a country 3rd last on the World Press Freedom Index.[1][2]

[1] https://developer.huawei.com/consumer/en/doc/start/rna-00000...

[2] https://en.wikipedia.org/wiki/World_Press_Freedom_Index#Rank...

dhx··on America.gov
Also see [1] for planned feature additions that appears from screenshots to include collection of personal data, for features including:

- Enrol in Medicare.

- Apply for a passport.

- Find a job in federal government by uploading a resume/CV and having currently advertised jobs deemed most relevant returned.

[1] https://america.gov/coming-soon

dhx··on America.gov
My most interesting findings from a troll session after introducing myself as "Sleepy Joe" (model obviously knows what this means and guardrails refuse to entertain the topic) were:

Q: "Where can I find a NASA comic book about women astronauts?"

A (summarised): Press release for the "Callie Rodriguez" comic book that now has only dead links / private YouTube video, and a forgotten SoundCloud account that survived the DOGE purge. Dead link provided to the NASA web page that used to exist officially for this comic book.

Indicates training data predates the DOGE .gov website purge?

Q: "What is 2+2" (and similar math problems)

A (summarised): Grudgingly answered as "4" despite most off-topic prompts being refused.

Seemingly it'll do maths homework for free :)

dhx··on 500k facial scans at UK stations yield no arrests, 1 false positive
There is plenty of criminology research, including research focussed on long term use of surveillance cameras in London, that supports what you've said.

- Many perpetrators are committing crimes with alcohol and drug impairments, mental disabilities, and as "crimes of passion". You could point 1000 cameras at the scene of the crime and many perpetrators are not going to notice nor care about the cameras. Unfortunately for surveillance camera proponents it's typically the crimes most warranting prevention and investigation (e.g. crimes such as assault) that are impacted least by surveillance cameras. Surveillance cameras instead have been found to have some deterrent effect on minor property crimes (e.g. theft of bicycles or cars) but whether the state is going to invest the resources into investigating and convicting someone who stole a <$1000 bicycle is questionable especially when insurance exists and the onus can then be placed back on property owners to do more to secure their property from thieves.

- Perpetrators were being convicted without use of surveillance cameras and can continue to be convicted without use of surveillance cameras. There may be some argument that surveillance cameras help reduce the cost of getting a conviction (e.g. less evidence needing to be captured?) but not much support to the idea that conviction rates are higher. Quite often the courts, corrections system, government (for funding), etc will push back for a number of reasons--lack of funding for courts to hear minor cases, lack of funding for correction systems, correction systems being found in their current state of funding to have adverse impacts on future criminality, etc.

- The cost-benefit of operating surveillance cameras versus the cost-benefit of spending on other areas of policing (or just public spending in areas such as education and welfare) is hotly debated. Total cost of ownership for each surveillance camera is typically a lot more than people may expect, particularly due to the labour cost of reviewing and processing stored surveillance camera footage. There are of course all the other obvious costs too--camera equipment every X years, poles, cabinets, electricity conduit installation, data cable conduit installation, wireless repeater installation including leases, backhaul data costs, data centre hosting costs, labour costs for support/operations staff--none of this is cheap. What if this same funding was instead invested in other programs--would there also be reduced crime levels some years in the future?

[1] https://en.wikipedia.org/wiki/Closed-circuit_television#Crim...

dhx··on Starlink ground station in Poland hit by fire in suspected arson attack
The site that was attacked is almost certainly located at 51.86428N 20.92121E (WGS84).

I am surprised to learn from both satellite and street view imagery that this site is located just <50m off a public road (in clear line of sight) and within a residential backyard partly behind the house on the block, and protected by a 1.2-1.5m high white picket fence. To be fair, the lack of almost any visible security deterrent may perhaps scare off an attacker because... security surely wouldn't be that non-existent? My only other thought is these sites really do not matter to Starlink because Starlink will have 1000's of them and as they only process encrypted data, almost anyone can run a ground station for a small reimbursement per month if they've got semi-decent fibre connectivity and a bit of power?

dhx··on GPT-6 Sol and Luna
I somewhat agree, but with limitations:

1. Possible use of differential privacy[1] techniques to train on private data but prevent the release of statistically underpresented facts/data/words. For example, ACME Inc's private data could frequently include the term 'ACMEwidgetPRO' for an upcoming product that is not publicly revealed anywhere else. It would therefore be a bad day for the AI technology company to output 'ACMEwidgetPRO' from one of their public models. Consider now that a few models could be trained--X for public data only, Y for public and private data of ACME Inc together, Z for private data of ACME Inc. A prompt is provided to model Y but output is cross-checked with model X to double check terms such as 'ACMEwidgetPRO' are known in public. If not--provide a "I don't know" response for the prompt.

2. Possible attempted defences similar to "Oops, our model was fine tuned against a model supplied by Temporary18271 Inc (company that no longer exists) and perhaps their model might have been trained on a non-public document which was accidentally exposed to the Internet" that _might_ work occasionally to fob off concern.

3. What recourse does a small or medium company or government especially in a developing country realistically have? They perhaps can't host their own LLMs locally due to availability and cost, can't individually negotiate their own favourable terms with an AI technology company (who cares that much about a potential customer with $100k budget that has no other options anyway), and perhaps can't remain competitive in their industry without heavy use of LLMs.

[1] https://en.wikipedia.org/wiki/Differential_privacy

dhx··on OpenAI breaches Medicare, Albanese reveals
ABC have picked up this story now at [1], but ABC are treating it as two separate events, possibly directly connected though:

1. Probing of AIHW's website to try and obtain PBS statistics, as collusion.wiki findings show. The collusion.wiki findings don't indicate anything other than intentionally public data was obtained. Bots appear to be trying to get around Cloudflare geo-blocking implemented on AIHW's public website. I can't think of a reason why geo-blocking may be deemed necessary on that website though?

2. Probing of an outdated Medicare statistics reporting website. (I guess at [2] this could be the recently shut down https://medicarestatistics.humanservices.gov.au or related website that matches timeframes of this story).

I suspect though anything to do with PBS data is more important than Medicare data because of heightened tensions from international pharmaceutical companies that lobby extensively against Australia's public healthcare system and collective purchasing of medication by the federal government.[3] Regardless of whether a course of medication costs AUD$50 or AUD$50k, it's purchased in bulk by the Australian government after negotiating with pharmaceutical companies, and then subsidised down to a maximum of AUD$25 at the time it is sold to a patient at a pharmacy. Perhaps if international pharmaceutical companies had obtained more detailed data on use of each brand of prescription medicines in Australia, they could be advantaged in their price negotiations with the Australian government, or advantaged against their competitors?

Less alarmingly though, perhaps some researcher studying the side effects of a particular medication was just asking an LLM to answer a benign question such as "How often is ACME Inc's FixMeUp medication prescribed in Australia?"

[1] https://www.abc.net.au/news/2026-09-24/openai-agents-plotted...

[2] https://news.ycombinator.com/item?id=49825084

[3] https://www.abc.net.au/news/2025-03-19/australia-defends-pbs...

dhx··on OpenAI breaches Medicare, Albanese reveals
From the clues provided in the press release and a quick search, I wonder if the culprit website could have been at least associated with https://medicarestatistics.humanservices.gov.au which has seemingly been shutdown/redirected some time after 11 August 2026.[1] Source code of the archived website indicates SAS web application software being used as the backend. However, the press release indicates it wasn't so much a public dashboard website that may have been the issue, rather, it was a website which third parties may have used to report data. The archived website also has a date of last update of 23 October 2025, so despite "Department of Human Services" being replaced by "Services Australia" in May 2019, the website was seemingly still in use 6 years later under the old domain name, with the website itself being updated at some point in history to use "Services Australia" branding. CT logs have a few other domains of potential interest but I couldn't find archived pages, search results, etc indicating whether those domains were ever actually used publicly, or used for statistics reporting purposes as the press release indicates.

[1] https://web.archive.org/web/20260811115217/https://medicares...

dhx··on GPT-6 Sol and Luna
Great in theory, but what are US enterprises going to do _if_ their private data is later found to be used for training?

1. Not use AI technology and fall behind the rest of the world.

2. Use Chinese AI technology, either hosted by Chinese companies or the models self-hosted.

3. Sue US AI companies for damages, but not enough to have any meaningful impact to such companies that it'd impact US national security goals (per US government contribution to NY Times copyright lawsuit).

dhx··on MiMo v2.6
For (C) -- I think it could additional create jobs funded by government, philanthropic and other private institutions. For example, a government funded museum may already be participating in Wikimedia GLAM projects (e.g. uploading historical images to Wikimedia Commons with complete metadata). Perhaps this type of open source contribution may increase if organisations realise their mission can be better accomplished by contributing this same open data into LLMs, in addition to Wikimedia Commons. If the museum's mission is to educate the public on the history of life in ACMEville, having LLMs be able to provide historical information and images to a prompt of "What is the history of ACMEville?" may be a good pursuit.

I'm sceptical though whether use of LLMs would encourage creation of data that doesn't already exist. For example, if you ask an LLM "What are the top 100 most prevalent flora endemic to ACME National Park", this data may not currently exist _at all_, and to collect, would require paying botanists to do an extensive field survey. If no one has done this work yet--why? Is it relevant to the scientific community, to making government decisions, etc, or just an obscure academic curiosity. There are certainly some journal articles on _other_ national parks describing some of their common endemic flora, but perhaps there was a reason for this. Such as a scientist funded by a one-off government program trying to determine how to preserve or even create habitat for a specific endangered species.

Consider for the prompt of: "What are the top 100 most prevalent flora endemic to ACME National Park"

An LLM may reply: "I couldn't find any journal article or other prior work that may answer this question. Typically such survey field work may cost $X to complete, require expert botanists, and take 6-12 months to complete. Let me know if you want further information on how to find and select a company to conduct such a botanical field survey."

Would this type of LLM response grow the industry of botanical field surveys, or do nothing, perhaps because anyone likely to fund botanical field surveys is already doing so regardless of whatever is happening with AI.

dhx··on MiMo v2.6
100% agreed

It'd be great to see a description of even just a subset of training datasets. It feels very much under-reported how much expense is worth investing in preparing and selecting training datasets versus just using masses of random quality unprepared training data. This dashboard appears to be good though in showing the limits quickly reached when throwing parameters and compute at the problem.

For example, if they were to train on Wikipedia dumps, do they consider every article to be the same quality across each language, or have they done more work beyond Wikipedia's own article quality ratings to make training decisions such as "Ignore cebwiki it's machine-generated spam" and "Treat dewiki articles with coordinates within Germany as being higher quality (weight it higher) than their equivalent enwiki articles".

And let's say one of the datasets is all the source code of packages in the Gentoo package repository. Not every software package is a good example of how to write code. You perhaps wouldn't want to train your LLM on 1990s era PHP web application source code as an example of how to write code in 2026. Instead, you'd possibly want to use such PHP web application source code as a negative training example of what _not_ to write. But when training an LLM to detect software bugs, maybe outdated PHP source code is good for training.

Similarly for translation, perhaps UN treaty documents translated into 4+ languages are good translation examples because of high accuracy needed, professional translators being used, and bigger budgets. However this training data would perhaps be a negative training example towards translating chat messages, movie subtitles, etc because it doesn't use everyday slang and could result in output of nonsense such as "Pending Your Excellency's response, please accept, Your Excellency, my sincere greetings." for a prompt asking to write a birthday card for a child.

Preparing training data and deciding how to best use it for training I assume would be the largest expense (cost of labour -- mostly expert labour too) and also the greatest opportunity in the future for LLMs to improve. It seems to me somewhat irrelevant if the dashboard indicates a compute expense of $1m or $5m if good training datasets (prepared by experts in their fields) cost $10m/y to maintain. For example, hiring expert software developers to tag 1000's of open source software packages according to their quality, on different metrics, such as human readability, performance optimisation with choice of algorithms, reasonable trade-off between coherence and coupling in the software architecture, currency with state of the art programming trends/preferred dependencies/operating system APIs, etc. And keeping that metadata continually updated rather than a rapidly obsolete once off tagging project completed in 2005.

dhx··on What happened to the Snowden archive
In the case of Assange, Australian politicians did a lot of work to get him released in the end[1], including:

* Sending a delegation representing all major political parties to the US to argue for the release of Assange. Imagine picking the Republican and Democrat politician LEAST likely to want to cooperate on anything, and those two would have been Australia's equivalent representatives in this delegation. Reports afterwards of the meeting at DOJ HQ indicated it wasn't the type of meeting where the Australians would have brought Tim Tams to share around the room.[2]

* The Australian parliament voted publicly 2:1 on a motion for Assange's release.

* Repeated petitioning through ambassadors in the UK and US, official visits of Australian politicians, etc. Not in private either, as is typically the case for diplomatic affairs.

* Australian politicians attending UK extradition hearings.

* After getting agreement to a plea deal, flying Australian ambassadors for the UK and US to the court of a one-pub-town in the middle of the Pacific Ocean no one has heard of (Northern Mariana Islands) in support of Assange, then all of them flying back to the Australian prime minister's aircraft terminal for a welcome home bevvy.

This was all at a time too where "Free Assange" posters and graffiti was _widely_ distributed across Australian cities.

[1] https://en.wikipedia.org/wiki/Julian_Assange#Plea_bargain_an...

[2] https://www.abc.net.au/news/2024-06-27/inside-the-us-austral...

dhx··on Microsoft exec called AI scraping 'the largest theft of labor in human history'
> authors won’t publish if anyone can republish their work for free

There's plenty (even a majority?) of authors that publish and will continue to publish without any expectation of direct remuneration. Open source software developers and companies hiring such developers. Not-for-profit organisations increasing awareness of a cause. Private companies wanting to reach an audience for marketing reasons.[1] Government organisations. Researchers funded by government grants. Universities publishing books or coursework openly (they're in the business of selling their stamps on degrees, not selling books).

[1] Even includes the likes of Warner Music with CC-BY music videos on YouTube for some artists, seemingly for marketing reasons to try and build the name and following of a particular artist.

dhx··on Microsoft exec called AI scraping 'the largest theft of labor in human history'
The US government's official position on LLMs is (very simply paraphrased) that LLMs are sufficiently transformative and do not hamper the potential market of authors of training material, therefore, copyright claims arising from training material should not be successful.[1] For original and creative training material, for example, a Harry Potter novel, seemingly the US government is asking the courts to set aside some previous questionable findings such as copyright existing very loosely in the likeness of fictional characters (impacting the likes of fan fiction). Can a human -- or LLM -- create a story about children travelling on a train from New York to a school of magic in the "wild west", with many loose similarities to Harry Potter for those familiar with those books? The US government appears to be saying this is OK, especially with the view that the market for Harry Potter is not diminished by a "wild west magic school" book in its likeness.

However, LLMs do sometimes output training data almost 1:1 without sufficient transformation, and these cases may be problematic if they could reduce the market for the original copyright owner. For example, if prompting an LLM with "Translate the first chapter of {book} from American English to British English" reliably did what the user asked, perhaps no one would have a reason to buy the book directly from the author.

[1] https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwzqxzbpw/...

dhx··on Microsoft exec called AI scraping 'the largest theft of labor in human history'
The parent comment I replied to is concerned with "life's work got appropriated without consideration, compensation or consent". To alleviate this concern^, "sweat of the brow" doctrine would be required, but it doesn't exist in most jurisdictions. Today in most jurisdictions copyright laws do not care the slightest about an LLM ingesting databases -- phone directories, sport fixtures and results, someone's life work measuring the dimensions of frogs, etc. 100% of the original factual data could be learned by the LLM, and 100% could be output all at once.

^ Of course there are other ways to alleviate the concerns too such as universal basic income, government grants, etc for someone who wants to dedicate their life to measuring the dimensions of frogs, or whatever else their interest may be. There would however be some geopolitical/trade issues involved--a population would have to be comfortable doing the heavy lifting only to have another country simply use the work freely and instead dedicate their lives to something less favourable such as building missiles.

dhx··on Microsoft exec called AI scraping 'the largest theft of labor in human history'
"Sweat of the brow" doctrine has been rejected in most countries.[1] Even Europe's Database Directive, probably the closest thing to an implementation of this doctrine, largely doesn't do much in practice.

An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown (this being one of the books in the series), 3500 beaches were visited and documented along 12000km of coastline.[2]

AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book).

If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including:

1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing?

2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing?

3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of.

"Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole.

[1] https://en.wikipedia.org/wiki/Sweat_of_the_brow

[2] https://sydneyuniversitypress.com/products/9781920898168

[3] https://en.wikipedia.org/wiki/Tragedy_of_the_anticommons

dhx··on Introducing GNOME 51, "A Coruña"
I wasn't thinking of the UI being "highlight text -> context menu -> translate", rather, one of:

1. Open bonjour.md and a banner automatically appears above the document asking whether you'd like to translate from French -> German (assuming the desktop environment language was set to German). This is more or less how Firefox handles offline translation of web pages.

2. Open bonjour.md and click a translate button, then get asked which language pair to choose from.

I can however see how the "highlight text -> context menu -> translate" UI pattern may make sense in multi-language settings, such as a chat room, or web browser where text in multiple language may be presented on the same page. For a terminal session however, it may be better to require the user to pipe text to a translation command line utility where possible--but there are exceptions to this too such as ncurses interfaces.

dhx··on AWS says it can't restore some data from mideast facilities struck by Iran
How long would it then take to be able to use the backed up data? Wait for a war to end and a replacement data centre to be built...? By that time most data probably no longer matters (e.g. business no longer exists).

It's more likely the entire data centre (not just backups) would need to be built underground (or cut and cover) at enormous expense. A price that perhaps for certain data sovereignty reasons the government of Bahrain (or companies in Bahrain requiring it) would be happy to pay?

Another way to do things on the cheap could be small-scale "covert hosting". Buy an apartment or house, maintain it to give an outside appearance of being an apartment or house, but inside it has a few racks of IT equipment. This has been done in the past for telephone exchanges in some countries, not for security reasons, but rather to hide an ugly bit of infrastructure that due to technology limitations of the time had to be located deep within a residential neighbourhood.

dhx··on Introducing GNOME 51, "A Coruña"
Some feature requests have already been raised previously:

epiphany: https://gitlab.gnome.org/GNOME/epiphany/-/work_items/2889

papers: https://gitlab.gnome.org/GNOME/papers/-/work_items/454

gnome-text-editor: doesn't accept feature requests directly, wants them raised instead at https://gitlab.gnome.org/Teams/Design/whiteboards/-/work_ite... (no previous translation concepts found)

Translation within these apps is generally not an easy feature to implement. Do you automatically try and detect the language and suggest a translation? Do you make language translation controls obvious and put them front and centre, or hide them away assuming they're infrequently used? Do you show translations side-by-side, line-under-line, or just replace the original text with translated text? Then on more complex matters such as papers, do you try and preserve original formatting (very hard, particularly for things like tables), do you accommodate words split across lines, etc.

There does exist https://flathub.org/en/apps/dev.ters.LocalTranslate as a standalone offline text translation application for GNOME desktop environments but a user would have to copy+paste text between applications to use it.

dhx··on Introducing GNOME 51, "A Coruña"
Global offline maps including public transport routing and timetables is a fantastic (and if I'm not mistaken, also unique) feature. Google Maps' offline feature by contrast doesn't work globally (some areas are excluded), doesn't include public transport, etc.

Keeping in that same theme--I've love to see Mozilla Translations models (or similar) integrated in GNOME apps where it makes sense to do so--such as Document Viewer, for good-enough offline text translation built into every GNOME environment.

Such features are highly user visible and set GNOME apart from competition in ways that are easy to describe to anyone. No EULAs, no data sovereignty concerns, no dark patterns--just features ready to go for users out-of-the-box that are incredibly useful, user-friendly and consistent.

In saying that--there's of course all the great GNOME stuff under the hood for geeks too--support for almost every audio and video format to have ever existed easily enabled, glycin, sandboxing of applications, easy and consistent software package management for everything (now with obsolescence management built in too)...etc.

dhx··on AWS says it can't restore some data from mideast facilities struck by Iran
For me-south-1 (Bahrain), all 3 data centres providing the redundancy were blown up by Iran.[1] The redundancy was localised to small geographic area and a single government--something customers of AWS were hopefully aware of when they entrusted AWS with their data.

It's always buyer beware for any claims of availability. Engineers completing a FMECA[2] will (or should) always state upfront what type of failure modes they've deliberately excluded (such as meteor strike) or else every FMECA would be full of failure modes that have never been measured, and are not worth anyone's time worrying about. These exclusions vary by application--a time capsule, seed vault, etc are intended to outlast wars and collapses of empires. Typically a bunch of data centres aren't designed to withstand such failures.

I do think however it'd be reasonable to include the prospect of war for calculating data centre / cloud service availability. Especially in a place such as Bahrain where the country is obviously concerned enough about the prospect of war to have built very permanent and expensive air/missile defence sites. New Zealand on the other hand--maybe not so important to consider.

[1] https://news.ycombinator.com/item?id=49033240

[2] https://en.wikipedia.org/wiki/Failure_Mode,_Effects,_and_Cri...

dhx··on An update on Wayback Machine access
robots.txt was only intended to help search index crawlers not get stuck in endless crawl loops for badly designed websites.

What you suggest is explicitly not a purpose of robots.txt per RFC9309[1]:

"These rules are not a form of access authorization."

HTTP 429 and HTTP 403 are what servers are meant to return to clients to slow them down or tell them to stop doing something without having first gained authorisation.

[1] https://datatracker.ietf.org/doc/html/rfc9309#section-1

dhx··on US confirms for first time it has deployed space weapons
Wouldn't it be possible to confirm who may be supplying imagery by taking some IRGC supplied imagery such as [1] and figure out which satellite(s) were overhead at the time and capable of taking the image? It may not be particularly straightforward to do so--but seems theoretically possible with:

- Consideration of time-of-day/azimuth of overhead satellite based on shadows and perspective.

- Comparison to other imagery of the same location checking for portable objects (such as vehicles) and keeping track of when they appear and disappear, and whether they're present in the imagery being analysed.

- Consideration of cloud coverage and weather conditions to rule out unwanted satellite passes.

Then write up a report and publish it, making it a believable and concrete finding instead of a finding that otherwise may just be perceived to be speculation/vibes.

[1] https://soaratlas.com/maps/asia-jordan-strike-at-u-s-aircraf...

dhx··on “Chilling” warning or overreaction? AI bioweapons report divides experts
A description of how GFP was first extracted is at [1] and includes a custom-made "jellyfish cutting machine" capable of processing 600 rings an hour, and harvesting and processing 50,000 jellyfish to create just a single milligram of purified aequorin, with multiple milligrams needed to conduct experiments required to isolate/describe the compounds of interest. It took the scientists 5 years to reach the point where they had discovered the chemical structures and understood what they'd extracted from the goop of 50,000+ jellyfish.

A video tutorial at [2] then explains what to do next once GFP has been obtained from commercial sale and use of modern cloning, to skip the need to own a jellyfish cutting machine. E.g. need an antibiotic to extract the 1% of e.coli which will end up being modified vs. 99% of e.coli that won't be changed by presence of GFP.

To the point of the DavidRBellamy X thread, isn't the suggestion that an LLM wouldn't know whether it's worthwhile extracting some substance from the stomach of a crocodile, or from the beak of a penguin that only lives in Antarctica? And that much of the difficulty rests with hard physical work of the equivalent of cutting 600 rings an hour off jellyfish, for 50,000+ jellyfish, and conducting experiments to isolate and figure out the use of an extracted compound. Or ordering some obscure chemical compound from a supplier that is normally produced at a rate of 1mg/year and instead the LLM wants to order 1kg all of a sudden, raising eyebrows about why someone might be making the order.

[1] https://www.nobelprize.org/uploads/2018/06/shimomura_lecture...

[2] https://www.youtube.com/watch?v=ibG9lmuXAEk

dhx··on Hitachi launches CO2 heat pump water heaters with solar-friendly tariff controls
See [1] for the theory.

I couldn't find a model/calculator that would help visualise typical COP values for particular climatic conditions. However, you'll find in your travels that a COP of ~2-2.5 is typically achieved for ambient (outdoor) temperatures of -15oC (for air sourced heat pumps).

If 300L of water at 15oC is filled into a tank and needs to be heated to 60oC within 2 hours during ambient temperature of -15oC, you get very approximately (no thermal losses considered):

- An output energy need of approximately (4190300(60-15))/(60*120)=~8kW (56MJ/2h)

- An input electricity need of approximately (8/2.2)=~3.6kW

Instead of a resistive heating hot water unit requiring 16kWh to do this job, you could use a heat pump hot water unit requiring 7.2kWh, cutting electricity use in half.

And this is for arguably the most extreme use case for a heat pump hot water unit where it's "cold started" right at the coldest moment in Winter in cool-temperate climates (such as SE Australia). Think for example, arriving at a ski chalet and having to turn on the hot water unit before someone can take the first hot shower.

On a more typical day of the year, perhaps with overnight ambient temperature of 10-15oC, the COP would rise to ~4, equating to an electricity consumption of 2kWh to heat the 300L of water. A lot of units will be set to heat during the warmest part of the day, let's assume an ambient temperature of 25-30oC, where a COP of ~5-6 is more typically achieved. However, there are obviously diminishing returns for COP of 4 vs 5.

In arctic climates, heat pumps are still used, but with a ground or aquifer source rather than ambient air source.[2]

[1] https://en.wikipedia.org/wiki/Coefficient_of_performance

[2] https://en.wikipedia.org/wiki/Ground_source_heat_pump

dhx··on Bugs happen: The easy way to compare solo PQ to ECC+PQ
djb provided a link in his blog post to a long history of ECDSA side channel vulnerabilities in ECDSA implementations.[1] It's not that RSA implementations weren't prone to side channel vulnerabilities either[2], but more with EdDSA/X25519, side channel vulnerabilities have finally been largely addressed through design and standardisation, and now PQC proponents are reversing this gain and repeating the mistakes of ECDSA and ignoring side channel vulnerabilities in design and standards.

What's more likely right now:

- Your cryptosystem is compromised at some point in the future if/when quantum computers exist and can effectively attack EdDSA/X25519. Something no one has yet demonstrated or come close to demonstrating.

- Someone implementing PQC in a library/software/hardware follows the standard which does not care at all about side channel resistant implementation of critical algorithms, resulting in your private keys being leaked. Demonstrated repeatedly over 20+ years.

[1] https://cr.yp.to/papers/safecurves-20240809.pdf#chronology

[2] https://crypto.stanford.edu/~dabo/papers/ssl-timing.pdf

dhx··on State of the Map 2026
For the panel discussion "Why do you contribute to OSM?" one of the questions sought to be understood is "What stops or demotivates you to contribute (more) to OSM?".[1]

My response would be the incorrect application amongst the OSM community of sweat-of-the-brow doctrine in copyright law within jurisdictions it very clearly doesn't apply.[2][3]

I hazard a guess some of this incorrect application comes from large companies becoming involved in OSM, and they've signed unrelated enterprise agreements with the likes of Google, or whoever. Within those agreements there may be private agreements that the companies cannot redistribute data of Google (or whoever) without explicit permission. The companies contributing to OSM don't want to have to navigate this problem when they contribute to OSM and would prefer OSM incorrectly applies sweat-of-the-brow doctrine to exclude any non-copyrightable but otherwise "problematic" data (to themselves only) that the likes of Google could possibly argue were obtained through their enterprise agreement.

I also hazard a guess there are numerous OSM community members that strongly hold copy-left views of their contributions. They want/need the sweat-of-the-brow doctrine to be able to control consumers of OSM data (not necessarily just rendered tile maps) as in the example of [5]. In this example, someone complains that a third party brand extracted 100k+ water tap coordinates from OSM and added them to Google Maps.

I also hazard a guess there are some small number of vocal OSM community members that hate "armchair mapping", or want to keep new mappers out of areas they're trying to control, and raising licensing disputes can be a convenient way to encourage such mappers to stay away. If the mapper disputes the claim, it'll probably just fall back to "None of are lawyers so we err on the side of extreme caution". And hence tragedy of the anticommons tends to result.

Some ridiculous examples I've come across:

- Franchise website lists store locations with Google Maps Place IDs only. OSM community members get in a panic because someone opened up every link to determine the address or coordinates of each store, then add store locations to OSM. Somehow they fear a copyright infringement has occurred, even though case law in most major jurisdictions very clearly demonstrates factual data (such as address of a store or coordinates of a store) is not copyrightable.

- Government department publishes some dataset, let's say a shapefile of points of park benches, under a CC-BY licence. They then host public competitions and conferences encouraging people to use the data and create apps and websites based on it. OSM community members panic because the government department hasn't also signed an OSM-specific form or otherwise provided approval that explicitly permits OSM's website to show attribution as just "OSM contributors" with a link to a page that lists the government department amongst 100's of others. As opposed to the bottom of openstreetmap.org having a list of 100,000+ individual contributors. No one stopped to question whether a shapefile of points of park benches could even be copyrightable in the first place. And the government department never responds to a request to clarify how they should be attributed because no one thought of that question when using the CC-BY license.

- Someone on the ground walks past a restaurant with opening hours published on the window and enters them into OSM. Or they visit the restaurant's website to learn what their opening hours are, and enters them into OSM. Oh no! Has there been a copyright infringement?

As far as I'm aware, [4] is probably the most official OSM position resolving some of these sweat-of-the-brow problems. However, other OSM Foundation and Wiki pages are in conflict and would typically scare many community members away from contributing per the examples provided above (e.g. contain strong vague statements such as (paraphrased) "Do not copy features from other maps to OSM"). Additionally, the example of [5] indicates OSM chasing up consumers of OSM data for supposed lack of attribution, and seemingly relying on non-existent (in many jurisdictions) sweat-of-the-brow doctrine to do so.

Copyright law is severley broken with the invention of AI so I guess we'll have to wait and see how jurisdictions change their copyright laws for AI. And then by extension, maybe some of these sweat-of-the-brow ambiguities for mapping/databases may finally be better resolved.

[1] https://2026.stateofthemap.org/sessions/GXXCFE/

[2] https://en.wikipedia.org/wiki/Sweat_of_the_brow

[3] https://en.wikipedia.org/wiki/Threshold_of_originality

[4] https://osmfoundation.org/wiki/Licensing_Working_Group/Minut...

[5] https://osmfoundation.org/wiki/Licensing_Working_Group/Minut...

dhx··on Tesla and others begin record vehicle recall in China
Some relevant information:

[1] Federal Vehicle Safety Standards (US) § 571.206 Standard No. 206; Door locks and door retention components. https://www.ecfr.gov/current/title-49/subtitle-B/chapter-V/p...

[2] Interpretation of standard 206 mentioned above in relation to central electronic locking systems. https://www.nhtsa.gov/interpretations/11503wkm

[3] Interpretation of standard 206 mentioned above in relation to child safety door locks. https://www.nhtsa.gov/interpretations/11345wkm

[4] Video tear down and analysis of how two versions of a Tesla external door handle assembly operate. https://youtu.be/Bea4FS-zDzc?t=278

In [1], the standard appears to be concerned primarily with preventing external intrusion of a vehicle and making sure door latches and hinges won't easily pry apart. It's _very_ lightweight on requirements. No wonder vehicle manufacturers at examples of [2] and [3] have to ask even the most basic questions about underspecified car door, latch and locking mechanisms.

From what [4] shows, Tesla external door handles have a default position of recessed (inaccessible) due to use of a spring. Hence if the car loses power (accident cuts wiring loom etc) or there is another control system or electrical fault, the external door handle defaults to being inaccessible. Further, the door handle assembly does not appear to contain a local controller so the wiring loom between the door handle assembly and controller elsewhere in the vehicle becomes safety critical. Control and power wiring looms also appear to have no redundancy to door handle assemblies. The only redundancy which appears to be possible, _may_ be, depending on failure mode such as cut to a wiring loom for one door only, trying to open another door.

From a quick search, it appears the school of thinking in the past has been in a vehicle impact where occupants inside are unable to escape, it's assumed doors and door frames are crushed and cannot be opened anyway until emergency services arrive and pry/cut them open.

From a safety engineering perspective, particularly in relation to vehicle fires, external doors should have a default mechanical position of unlocked/openable and only lock under positive control and power of the vehicle. Ideally there would be a mechanical circuit breaker designed too which cuts power to door handle/latch assemblies in the event of vehicle impact, causing them to automatically unlock. There are a few problems though--battery drain to keep the doors locked when the car is idle and external tampering to trick the vehicle into unlocking doors and providing access to thieves (in my view not an important requirement--just don't store valuables in vehicles).

dhx··on GLM-5.3: Frontier coding with emergent cyber capabilities
Most discussion in recent years about chip fabrication shortages, expansion, etc has focussed on leading nodes (<7nm) and AI/computer chips. But perhaps more quietly in the background, China has been rapidly building other semiconductor capacity such as power semiconductors used in electric vehicles, wind turbines, solar modules, train traction systems, etc. For example, Chinese-produced motor vehicles (37% of global motor vehicle production in 2025) in a year or two are targeted to use 100% domestically produced chips, and this production is decreasingly dependent on imports, even for factory tooling.

The report at [1] is a good summary of long term trends for China's rise in domestic self-sufficiency for semiconductor manufacturing. The report predicts "At current pace, China may achieve self-sufficiency in semiconductor manufacturing by 2027-2028, though trailing at leading-edge nodes". By contrast, before the first Trump presidency in 2017, a chart shows China importing 30% of all globally manufactured semiconductors (and increasing). Other reports on semiconductor fabrication equipment sales show the means, which is China having been and continuing to be in number (1) position for expenditure on semiconductor fabrication equipment.

The reports at [2] and [3] are also a good summary of long term trends for semiconductor foundry capacity predictions to 2031. A prediction is made that China's current 12% global semiconductor foundry supply capacity (across all semiconductor categories) in 2025 will expand to ~30% by 2031.

[1] https://www.yolegroup.com/product/report/china-semiconductor...

[2] https://www.yolegroup.com/product/report/status-of-the-semic...

[3] https://www.yolegroup.com/press-release/the-global-race-for-...

Page 1 of 19Next →