HNHacker News
TopNewBestAskShowJobs

Deathmax

500 karma · joined August 10, 2012

submissionscomments
Deathmax··on Google Gemini has the worst LLM API
Can we avoid weekend changes to the API? I know it's all non-GA, but having `includeThoughts` suddenly work at ~10AM UTC on a Sunday and the raw thoughts being returned after they were removed is nice, but disruptive.
Deathmax··on Google Gemini has the worst LLM API
> You can’t just attach an image to your request.

You can? Google limits HTTP requests to 20MB, but both the Gemini API and Vertex AI API support embedded base64-encoded files and public URLs. The Gemini API supports attaching files that are uploaded to their Files API, and the Vertex AI API supports files uploaded to Google Cloud Storage.

Deathmax··on Gemini 2.5 Flash
It is gated behind the GOOGLE_INTERNAL visibility flag, which only internal Google projects and Cursor have at the moment as far as I know.
Deathmax··on How to gain code execution on hundreds of millions of people and popular apps
That limitation should go away when Trusted Signing graduates from preview to GA. The current limitation is because the CA rules say you must perform identity validation of the requester for orgs younger than 3 years old, which Microsoft isn't set up for yet.
Deathmax··on Diablo hackers uncovered a speedrun scandal
Their primary business model nowadays is as an advertising agency, not book selling: https://www.guinnessworldrecords.com/business-marketing-solu...
Deathmax··on Nvidia's RTX 5090 power connectors are melting
The cables between 12V-2x6 and 12VHPWR are identical, it's the port that has different pin lengths (shorter sensing pin, longer conductor pins) to allow for better detection of poorly seated cables and better conductivity while loose.
Deathmax··on VSCode’s SSH agent is bananas
It's helpful for evading detection, because if you've compromised a machine, you can drop in the server binary and it'll have been added to the allowlist for devs to run.
Deathmax··on Llama.cpp supports Vulkan. why doesn't Ollama?
It's simple enough to test the tokenizer to determine the base model in use (DeepSeek V3, or a Llama 3/Qwen 2.5 distill).

Using the text "സ്മാർട്ട്", Qwen 2.5 tokenizes as 10 tokens, Llama 3 as 13, and DeepSeek V3 as 8.

Using DeepSeek's chat frontend, both DeepSeek V3 and R1 returns the following response (SSE events edited for brevity):

  {"content":"സ","type":"text"},"chunk_token_usage":1
  {"content":"്മ","type":"text"},"chunk_token_usage":2
  {"content":"ാ","type":"text"},"chunk_token_usage":1
  {"content":"ർ","type":"text"},"chunk_token_usage":1
  {"content":"ട","type":"text"},"chunk_token_usage":1
  {"content":"്ട","type":"text"},"chunk_token_usage":1
  {"content":"്","type":"text"},"chunk_token_usage":1
which totals to 8, as expected for DeepSeek V3's tokenizer.
Deathmax··on Llama.cpp supports Vulkan. why doesn't Ollama?
The most recent one of the top of my head is their horrendous aliasing of DeepSeek R1 on their model hub, misleading users into thinking they are running the full model but really anything but the 671b alias is one of the distilled models. This has already led to lots of people claiming that they are running R1 locally when they are not.
Deathmax··on We shrunk our Javascript monorepo git size
According to ThinkBroadband's tracking [1], the headline figures are 85.20% of premises are gigabit capable (FTTP/FTTH/Cable [DOCSIS]) with 71.86% being full fibre.

[1]: https://www.thinkbroadband.com/news/10343-85-gigabit-coverag...

Deathmax··on Show HN: Detect if an audio file was generated by NotebookLM
As far as I know, there's no available tooling for the public to detect SynthID watermarks on generated text, image, or audio, outside of Google Search's About this Image feature.
Deathmax··on Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
The free tier API isn't US-only, Google has removed the free tier restriction for UK/EEA countries for a while now, with the added bonus of not training on your data if making a request from the UK/CH/EEA.
Deathmax··on Two new Gemini models, reduced 1.5 Pro pricing, increased rate limits, and more
They do: https://cloud.google.com/vertex-ai/generative-ai/docs/partne...
Deathmax··on Two new Gemini models, reduced 1.5 Pro pricing, increased rate limits, and more
They do mention how characters are counted in the Vertex AI pricing docs: "Characters are counted by UTF-8 code points and white space is excluded from the count"
Deathmax··on Two new Gemini models, reduced 1.5 Pro pricing, increased rate limits, and more
There's some possible confusion because of the Copilot problem where everything in the product stack is called Gemini.

The Gemini API (or Generative Language API) as documented on https://ai.google.dev uses https://ai.google.dev/gemini-api/terms for its terms. Paid usage, or usage from a UK/CH/EEA geolocated IP address will not be used for training.

Then there's Google Cloud's Vertex AI Generative AI offering, which has https://cloud.google.com/vertex-ai/generative-ai/docs/data-g.... Data is not used for training, and you can opt out of the 24 hour prompt cache to effectively be zero retention.

And then there's all the different consumer facing Gemini things. The chatbot at https://gemini.google.com/ (and the Gemini app) uses data for training by default: https://support.google.com/gemini/answer/13594961l, unless you pay for Gemini Enterprise as part of Gemini for Workspace.

Gemini in Chrome DevTools uses data for training (https://developer.chrome.com/docs/devtools/console/understan...).

Enterprise features like Gemini for Workspace (generative AI features in the office suite), Gemini for Google Cloud (generative AI features in GCP), Gemini Code Assist, Gemini in BigQuery/SecOps/etc do not use data for training.

Deathmax··on Qwen2.5: A Party of Foundation Models
They've posted their own run of the Aider benchmark [1] if you want to compare, it achieved 57.1%.

[1]: https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2.5/Qwen...

Deathmax··on Pixtral 12B
That hasn't been true for their largest model since the 2407 release of Mistral Large 2 (https://mistral.ai/news/mistral-large-2407/), it is however under a non-commercial license.
Deathmax··on What is an SBAT and why does everyone suddenly care
That would be compromising the domain owner, rather than the threat model of Certificate Transparency which is compromised Certificate Authorities, especially given the number of government owned, publicly trusted (sub-)CAs.
Deathmax··on ShadPS4 – PlayStation 4 emulator
Emulating a PS3 is significantly harder than emulating a PS4, since on the PS3 you need to convert from the Cell architecture, whereas the PS4 is already x86.
Deathmax··on Ask HN: Do we need to pay billions in fees to Stripe, Block, PayPal and Visa/MC?
> As long as debit cards have a magnetic stripe and have their full number printed on them, and that information is useful, this problem remains.

Which the EEA/UK has also (partially) solved by enforcing Strong Customer Authentication (SCA) that mandates that (most) transactions require MFA.

Deathmax··on 'Kafkaesque': bank blocks cash transfer, saying it could be an AI scam
As I understand it, UK banks (where this article is written) will soon be required[1] to reimburse victims of Authorised Push Payment fraud (some banks are signed up for a voluntary code at the moment) outside of "gross negligence" by the customer, so they can't just allow customers to waive liability. Making transfers now involves multi step questionnaires with big scary warnings about scams at every stage.

[1]: https://www.psr.org.uk/publications/policy-statements/ps233-...

Deathmax··on Show HN: Open-source LLM provider price comparison
For throughput data, well, you need to actually run prompts to gather the data which racks up costs fast and performance can vary based on input prompt lengths. The two sources I use are OpenRouter's provider breakdown [1] and Unify's runtime benchmarks [2].

[1]: https://openrouter.ai/models/meta-llama/llama-3.1-70b-instru...

[2]: https://unify.ai/benchmarks/llama-3.1-70b-chat

Deathmax··on Make your electronics tamper-evident
Showing how often an authenticity code has been checked is something manufacturers like Xiaomi do, where there's rampant counterfeiting.
Deathmax··on Threat actor abuses Cloudflare tunnels to deliver remote access trojans
The operative word being most registrars. If you look at the list of registrars commonly used by bad actors, you can find a list of registrars that are either non-responsive to abuse complaints, or only take action after n days.
Deathmax··on Researcher finds flaw in a16z website that exposed some company data
You don't want to push secrets in their raw form on GitHub, secret scanning would disable keys from supported providers.
Deathmax··on Gemma 2: Improving Open Language Models at a Practical Size [pdf]
Gemini models on Vertex AI can be called via a preview OpenAI-compatible endpoint [1], but shoving it into existing tooling where you don't have programmatic control over the API key and is long lived is non-trivial because GCP uses short lived access tokens (and long-lived ones are not great security-wise).

Billing for the Gemini models (on Vertex AI, the Generative Language AI variant still charges by tokens) I would argue is simpler than every other provider, simply because you're charged by characters/image/video-second/audio-second and don't need to run a tokenizer (if it's even available cough Claude 3 and Gemini) and having to figure out what the chat template is to calculate the token cost per message [2] or figure out how to calculate tokens for an image [3] to get cost estimates before actually submitting the request and getting usage info back.

[1]: https://cloud.google.com/vertex-ai/generative-ai/docs/multim...

[2]: https://platform.openai.com/docs/guides/text-generation/mana...

[3]: https://platform.openai.com/docs/guides/vision/calculating-c...

Deathmax··on Apple's On-Device and Server Foundation Models
On Apple’s side, Apple Intelligence will only be enabled on A17 Pro and M-series chips, so only the iPhone 15 Pro and Pro Max will be supported in terms of phones.
Deathmax··on What to do when an airline website doesn't accept your legal name
I would check what's encoded on the machine readable zone on your passport, as I understand it, the only valid characters are A-Z, 0-9, and < as a field separator.
Deathmax··on Gemini 1.5 outshines GPT-4-Turbo-128K on long code prompts, HVM author
https://support.google.com/gemini?p=activity_off_retention

"Who has access to my Gemini Apps conversations?

How you can control what’s shared with reviewers

If you turn off Gemini Apps Activity, future conversations won’t be sent for human review or used to improve our generative machine-learning models."

The way I interpret the policy is that if Gemini Apps Activity is off:

1. Conversations after activity is turned off that are linked to your account are retained for up to 72 hours for "safety and security", but are not sent to human reviewers or used for model training.

2. In the event that you submit feedback (eg good/bad response), then the conversation (disassociated from your account) can be used for training. This is also flagged when attempting to submit feedback: "Even when Gemini Apps activity is off, feedback submitted will also include up to the last 24 hours of your conversations to help improve Gemini."

Deathmax··on Gemini 1.5 outshines GPT-4-Turbo-128K on long code prompts, HVM author
The Google One AI plan does provide access to Gemini Advanced (backed by Gemini Ultra 1.0) at a significant premium to the current highest Google One plan: https://one.google.com/explore-plan/gemini-advanced

The free tier of Gemini (formerly Bard) has been on Gemini Pro 1.0 for a bit now.

← PreviousPage 2 of 5Next →