HNHacker News
TopNewBestAskShowJobs

msoad

8,467 karma · joined January 12, 2013

submissionscomments
msoad··on QwQ: Alibaba's O1-like reasoning LLM
no the OP but literally your comment as prompt

https://chatgpt.com/share/6747c7d9-47e8-8007-a174-f977ef82f5...

msoad··on QwQ: Alibaba's O1-like reasoning LLM
Somehow o1-preview did not find the answer to the example question. It hallucinated a wrong answer as correct. It eventually came up with another correct answer:

    (1 + 2) × 3 + 4 × 5 + (6 × 7 + 8) × 9 = 479

Source: https://chatgpt.com/share/6747c32e-1e60-8007-9361-26305101ce...
msoad··on Amazon S3 Adds Put-If-Match (Compare-and-Swap)
It ensures that when you try to upload (or “put”) a new version of a file, the operation only succeeds if the file on the server still has the exact version (ETag) you specify. If someone else has updated the file in the meantime, your upload is blocked to prevent overwriting their changes.

This is especially useful in scenarios where multiple users or processes are working on the same data, as it helps maintain consistency and avoids accidental overwrites.

This is using the same mechanism as HTTP's `If-None-Match` header so it's easier to implement/learn

msoad··on Amazon S3 now supports the ability to append data to an object
Does it really work for livestreams? Can I stream read and write on the same video file? That is huge if true!

Edit: oh it’s only in one AZ

msoad··on Monorepo – Our Experience
I love monorepos but I'm not sure if Git is the right tool beyond certain scale. Where I work doing a simple `git status` takes seconds due to the size of the repo. There has been various attempts to solve Git performance but so far this is nothing close to what I experienced at Google.

The Git team should really invest in tooling for very large repos. Our repo is around 10M files and 100M lines of code and no amount of hacks on top of Git (cache, sparse checkout etc etc) is not really solving the core problem.

Meta and Google have really solved this problem internally but there is no real open source solution that works for everyone out there.

msoad··on ChatGPT Search
Wow! This is the real news here!
msoad··on Launch HN: GPT Driver (YC S21) – End-to-end app testing in natural language
Where I work there are 1,500 microservices. How do I get a log of all of those services -- only related to my test's requests in a file?

I know there are solutions for this, but in the real world I have not seen it properly implemented.

msoad··on Launch HN: GPT Driver (YC S21) – End-to-end app testing in natural language
Exactly! I've never seen a 5000+ eng org that have all their ducks in a row when it comes to telemetry. it's one of those things that you can't put a team in charge of it and get results. everyone have to be on the same page which in a big org is hardly the case.
msoad··on Launch HN: GPT Driver (YC S21) – End-to-end app testing in natural language
I work in this space. We manage thousands of e2e tests. The pain has never been in writing the tests. Frameworks like Playwright are great at the UX. And having code editors like Cursor makes it even easier to write the tests. Now, if I could show Cursor the browser, it would be even better, but that doesn’t work today since most multimodal models are too slow to understand screenshots.

It used to be that the frontend was very fragile. XVFB, Selenium, ChromeDriver, etc., used to be the cause of pain, but recently the frontend frameworks and browser automation have been solid. Headless Chrome hardly lets us down.

The biggest pain in e2e testing is that tests fail for reasons that are hard to understand and debug. This is a very, very difficult thing to automate and requires AGI-level intelligence to really build a system that can go read the logs of some random service deep in our service mesh to understand why an e2e test fails. When an e2e test flakes, in a lot of cases we ignore it. I have been in other orgs where this is the case too. I wish there was a system that would follow up and generate a report that says, “This e2e test failed because service XYZ had a null pointer exception in this line,” but that doesn’t exist today. In most of the companies I’ve been at, we had complex enough infra that the error message never makes it to the frontend so we can see it in the logs. OpenTelemetry and other tools are promising, but again, I’ve never seen good enough infra that puts that all together.

Writing tests is not a pain point worth buying a solution for, in my case.

My 2c. Hopefully it’s helpful and not too cynical.

msoad··on Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
I skimmed through the computer use code. It's possible to build this with other AI providers too. For instance you can asks ChatGPT API to call functions for click and scroll and type with specific parameters and execute them using OS's APIs (A11y APIs usually)

Did I miss something? Did they have to make changes to the model for this?

msoad··on Grandmaster-level chess without search
Searched only once. If this can be applied to other knowledge with this efficiency we're onto something
msoad··on Upgrading Uber's MySQL Fleet
Yeah, I kinda stopped reading when I felt this. Not sure why? The substance is still interesting and worth learning from but knowing LLM wrote it made me feel icky a little bit
msoad··on Understanding the Limitations of Mathematical Reasoning in Large Language Models
A calculator can "think" is "AI" and however you want to frame it. Reasoning is a very specific and defined concept. Computers can not reason (per this paper)
msoad··on Deno 2
being able to run Next.js with Deno is huge! I will give this a try!
msoad··on Differential Transformer
Like most things in this new world of Machine Learning, I'm really confused why this works?

The analogy to noise-cancelling headphones is helpful but in that case we clearly know which is signal and which is noise. Here, if we knew why would we even bother to the noise-cancelling work?

msoad··on Founder Mode, hackers, and being bored by tech
exactly! tech needs more Woz+Jobs and less Cooks.
msoad··on YAML is not a superset of JSON
`'{"a": 1e2}'` is YAML not JSON. What a weird argument.
msoad··on SwiftUI for Mac 2024
This is similar situation with Copilot. it does not suggest more modern web and JS APIs for me.
msoad··on Examples of Great URL Design (2023)
That's not the question. The question is what if you have an open ended URL param that can also have subpaths?
msoad··on Examples of Great URL Design (2023)
One big question in URL design is this:

Do path parameters get to have / in their values?

Let’s say you have a link shortener service and want to allow users to define shortcuts like /mypath/:rest where rest is appended to example.com/

Now you’re in a very interesting position when it comes to resolving URLs.

Curious to hear folks with experience in this

msoad··on Examples of Great URL Design (2023)
The first example is just that. Put the id in the URL and make the slug optional.

Stackoverflow makes the slug completely optional but you have the choice of only accepting foo and bar in your example

msoad··on Launch HN: Stack Auth (YC S24) – An Open-Source Auth0/Clerk Alternative
Can I have my own user database table without setting up web hooks?
msoad··on RLHF is just barely RL
A program like

    function add(a,b) {
      return 4
    }
Passes the test
msoad··on Structured Outputs in the API
Why not JSON Schema?
msoad··on Meta Reports Second Quarter 2024 Results [pdf]
A 9% improvement in margins is pretty good but I'm wondering if the margin was lower due to fine or something like that in '23?
msoad··on AI models collapse when trained on recursively generated data
wow this is such a good point! Evolution is just that!
msoad··on Alexa is in millions of households and Amazon is losing billions
> The technology isn’t there, but they have a deadline

The technology is there, as demonstrated by OpenAI's ChatGPT Voice Mode. With the resources and talent that Amazon possesses, they should at least be able to demo something similar. It's just that the Alexa organization is a mess, which prevents it from happening.

msoad··on Llama 3.1: Our most capable models to date
MMLU PRO is the benchmark I trust the most. I noticed they are using 5 shots and CoT. Is that true for GPT4 and Sonnet as well?
msoad··on Never Update Anything
Kinda ironic that the article itself was updated
msoad··on CSS Classes Considered Harmful
As this article argues, problem with CSS Classes is that a CSS Class can be applied to _any_ element where custom properties will be scoped to the elements that the CSS was meant to be used for.
← PreviousPage 3 of 34Next →