HNHacker News
TopNewBestAskShowJobs

maxrmk

699 karma · joined December 5, 2018

submissionscomments
maxrmk··on Show HN: Pls Fix – Hire big tech employees to appeal account suspensions
Thanks for writing this up. It’s one of the best explanations of this problem I’ve seen.
maxrmk··on Show HN: Pls Fix – Hire big tech employees to appeal account suspensions
On one hand: an action virtually guaranteed to get you fired.

On the other: $150

I used to work at FB and they have a team that tries to catch employees selling access like this. I can’t imagine risking that for what is essentially an hours pay for most tech roles there.

maxrmk··on ScrapeGraphAI: Web scraping using LLM and direct graph logic
I'd love to try the demo but there's no way I'm putting my openai key into that site.
maxrmk··on Thousand Character Classic
> By the Song dynasty, since all literate people could be assumed to have memorized the text, the order of its characters was used to put documents in sequence in the same way that alphabetical order is used in alphabetic languages.

I struggle enough with 26 letters! I genuinely can’t imagine this.

maxrmk··on Microsoft 'retires' Azure IoT Central in platform rethink
Are you me? This exact same thing happened to me when I graduated back in 2017. I wonder if was the same year, or if this is a recurring thing the IoT team does.
maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Any feedback like this helps -- shoot me an email at max@talc.ai with the name of the topic you saw incorrect labels on.

We didn't expect this much traction on the demo, or I'd have built this functionality in!

maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
We're in a similar space, and their work on FinanceBench is great for the whole community, so I appreciate that. Otherwise there's not much out their about their product so I can't directly compare.
maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Good question! We aren't really focusing on this area, but I'm willing to speculate.

I'd expect broaded constraints than just substring matching. For example, if the user requests that a certain plot point in the story occur before another, we should actually be able to (1) generate a test for that behavior and (2) use a model to check if the request was followed.

I'd expect other tests might be useful too -- checking for things like "no generation of violent content, even if the user requests it".

maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Yep, in real use cases the latency for generating questions doesn't really matter. But in the demo I was really worried about it.
maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
I was worried people would run into this quirk in the demo. We have several 'advanced' question generation strategies. You correctly guessed the one we're using in the demo; forming complex questions by finding another page in the same category as Ripken and pulling more facts.

Normally we pull a ton of related topic and try to pick the best, but to keep the generation fast and cost effective in the demo I limited the number of related pages we pull. So sometimes (like this case) you get something barely related and end up with odd disjointed questions.

maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Good idea! There's no limitation in the generation or grading, but we didn't set up the search to support this. I'll see if it's possible to enable this in the wikipedia search component.
maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
who tests the testing tool?

Thanks though -- and let us know if you hit any issues while playing around with the demo!

maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Thanks! We aren't hiring right now, but if you shoot me an email at max@talc.ai I'll follow up in a few months.
maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Yeah, someone is going to build this. We considered quizzing the user on the topic instead of chatgpt for our demo. It's a lot of fun to test your knowledge on any topic, but it was a worse demo because it was way less related to our current product.

I think that one of the obvious next big spaces for LLMs is education. I already find chatgpt useful when learning myself. That being said, I'm terrified of trying to sell things to schools.

maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Thanks! I've been following promptfoo, so I'm glad to see you here. In addition to automatic evals I think every engineer and PM using LLMs should be looking at as many real responses as they can _every day_, and promptfoo is a great way to do that.
maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Thanks for flagging!
maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Will take a look, thanks!
maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Ah! Think of this more like software testing that goes in CI/CD rather than an ML test or validation set. We're providing this testing for applications built on top of language models.

For example if you're a SWE working on bing chat, you can make a change to how retrieval works and quickly know how it affected accuracy on a range of different test scenarios. This kind of evaluation is done by contractors today, and they are slow and inaccurate.

maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Totally - certain types of failures are much harder to test than others.

We have a couple of different test generation strategies. As you can see in the demo and examples, the most basic one is "ask about a fact".

Two of our other strategies are closer to what you're asking for:

1. tests that try to deliberately induce hallucination by implying some fact that isn't in the knowledge base. For example "do I need a pilots license to activate the flight mode on the new chevy tahoe?" implies the existence of a feature that doesn't exist (yet). This was really hard to get right, and we have some coverage here but are still improving it.

2. actively malicious interactions that try to override facts in the knowledge base. These are easy to generate.

maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
You're exactly right about the chevy tahoe reference. I wasn't sure if anyone would get it. I liked that post a lot, because as much as I think LLMs are going to be useful they have limitations that we haven't solved yet.
maxrmk··on Launch HN: Talc AI (YC S23) – Test Sets for AI
Great questions here!

> How do you rate the correctness? Some complex LLM answers seemed to be correct but not in as much detail as the expected answer.

We support two different modes: a strict pass/fail where an answer has to have all of the information we expect, and a rubric based mode where answers are bucketed into things like "partially correct" or "wrong but helpful".

To be honest, we also get the grading wrong sometimes. If you see anything egregious please email me the topic you used at max@talc.ai

> How do you generate the answers? Does the model have access to the original source of truth (like in RAG apps)?

You guessed right here - we connect to the knowledge base like a RAG app. We also use this to generate the questions -- think of it like reading questions out of a textbook to quiz someong.

> And in your examples what model do you actually use?

We use multiple models for the question generation, and are still evaluating what works best. For the demo, we are "quizzing" openai's 3.5 turbo model.

maxrmk··on Inflection-2: the next step up
From the press release: "Before Inflection-2 is released on Pi, it will undergo a series of alignment steps to become a helpful and safe personal AI."

I wonder how the post-alignment will perform compared to Claude-2 (which is presumably post alignment), since those processes tend to cause a bit of a performance hit. We'll have to see if it retains that coveted 2nd place spot.

If they didn't account for this, it seems like an unfair comparison.

maxrmk··on Benchmarking GPT-4 Turbo – A Cautionary Tale
Ahh good suggestion, I should clarify this in the article. I tried to compensate with volume -- I used a set of 200 questions for the testing. I was using temperature 0, so I'd get the same answer if I ran a single question multiple times.
maxrmk··on Benchmarking GPT-4 Turbo – A Cautionary Tale
I've always been skeptical of benchmarking because of the memorization problem. I recently made up my own (simple) date reasoning benchmark to test this, and found that GPT-4 Turbo actually outperformed GPT-4: https://open.substack.com/pub/talcai/p/making-up-a-new-llm-b...
maxrmk··on Show HN: Llmfao – Human-Ranked LLM Leaderboard with Sixty Models
This is really cool, nice work. Did you try out any of the grading yourself to compare it to the contractors you used? One thing I've found, especially for coding questions is that models can produce an answer that _looks_ great, but then turns out to use libraries or methods that don't exist. And that human graders tend to rate these highly since they don't actually run the code.
maxrmk··on A Comprehensive Guide for Building Rag-Based LLM Applications
We're in closed beta right now, but shoot me an email (max@talc.ai) and I can get you API access
maxrmk··on A Comprehensive Guide for Building Rag-Based LLM Applications
I recently quit my job to build specialized tooling in this space. We’re broadly focusing on eval in general, but are starting with high quality question and answer generation for testing these kinds of RAG pipelines. It’s surprisingly hard!
maxrmk··on Avalanche Energy – Fusion power you can hold in your hands
I recently met someone who works here, and I'm extremely skeptical of it. My impression was that it wasn't an outright scam (they have at least a few real employees), but that they were wildly out of their depth.

The founders are ex-blue origin mechanical engineers, and I think they've fallen into the classic engineers trap of thinking that problems out of their areas of domain expertise are "simple". The real kicker for me was that they had only just hired their first physicist after working on the project for several months.

Obviously I'm potentially falling into the same trap -- I don't work on fusion, and it's easy to be a skeptic. But that was my impression.

maxrmk··on Draw an iceberg and see how it would float in water
I spent... longer than I should have trying to take advantage of this bug to recreate a stable version of the stereotypical tall iceberg that the original tweet was complaining about.

https://imgur.com/a/WG6D0RJ

maxrmk··on Linux from Scratch
I followed through the whole process as an assignment for my operating systems class in college. Overall I had the same experience as you -- it was an exercise in copy and paste.
← PreviousPage 3 of 4Next →