HNHacker News
TopNewBestAskShowJobs

gmays

45,899 karma · joined November 11, 2012

Building AI stuff.

Blog: https://gmays.com

https://github.com/gmays

Habit tracker: https://maincharacter.game/u/gabe

X: https://x.com/gabemays

LinkedIn: https://www.linkedin.com/in/gabrielmays

submissionscomments
gmays··on Show HN: Make math automatic with Mathy
Yeah, common core doesn't really give you the graph at the granularity I'm using. I mostly use it as a curriculum/placement authority.

The more granular decomposition is something I built separately. I take a standard/topic and break it down into coherent concepts, then into the smallest independently useful retrieval skills I want to track. For Mathy that's especially constrained because I only care about things that make sense for automaticity, not everything you'd need to teach the subject.

So for example a standard might imply multiplication/division fluency, but internally that gets decomposed into individual mathematical relationships and retrieval directions. `7 × 8`, `56 ÷ 7`, and `56 ÷ 8` are related, but they're not necessarily the same learning state.

Then for procedural material I define bounded problem families rather than copying individual questions. A family specifies the operand domain, mathematical rule, intended strategy, answer form, exclusions, hardest cases, etc, and the questions are generated deterministically from that.

So unfortunately I don't have a magical curriculum-generation hack :) It's been a combination of standards/course outlines for scope + a lot of decomposition + AI helping propose/check things + deterministic validation afterward. It was a lot of experimentation, and I still have room to improve.

I also intentionally don't copy questions from textbooks. The standards and public course outlines tell me what belongs where, but the actual problems/content are independently authored or generated. That makes it much easier to reason about correctness, coverage, etc.

For the domaine, yea, it's important. I built-in steaks and XP/levels since I could borrow the general approach from a previous habit tracking app I made (maincharacter.game). I also added medals for achievements. The website version has a leaderboard, which I may also add to the iOS app, but would take some add'l work since it's currently offline only ATM.

gmays··on Show HN: Make math automatic with Mathy
Thanks for the suggestion. I've considered this, but since it's a free app I didn't want to expand the maintenance burden beyond what I can keep up with.

There are some other things I want to improve first, but I'll definitely consider this if I see more demand/requests for it.

gmays··on Show HN: Make math automatic with Mathy
Nope, I haven't compared those. It roughly follows common core, so it should be somewhat similar, but I'm not sure how each of those based their curriculum.

Plus I added automaticity related topics from some Math Academy courses I want to take to help raise my learning ceiling.

gmays··on Show HN: Make math automatic with Mathy
Thanks. You might be thinking of the order of operations wrong. Example:

−3² = −(3 × 3) = −9 (−3)² = (−3) × (−3) = 9

If there are no parentheses, you first square 3 and then negate the result. But if there are, then you square -3 first as you explained.

I took a look and Mathy correctly accounts for these. But if you're sure your problem had parentheses instead of what you shared, please let me know!

Either way, thank you for the feedback, this was a good thing for me to double check.

gmays··on Show HN: Make math automatic with Mathy
Thanks for the feedback.

The goal of the product is to give practice and help build automaticity in topics you already know, which makes learning harder stuff (e.g. through Math Academy) easier later to raise your ceiling.

So, I'm not sure if maybe you're less familiar with the material (below the threshold) and you should take a course focused on teaching the material OR that the material in Mathy is too hard for the selected topic. Because if you're familiar, it should be helping you rather than frustrating you.

It's an offline app so I don't capture that kind of data. But you make a good point, and at a minimum for each topic I can have it start with easier problems then increasingly get harder as you do better.

I'll think more about that today and maybe get a fix out once I have some time.

Thanks again for the feedback!

gmays··on Show HN: Make math automatic with Mathy
Thanks. I added some structure that made it easier for me to QA, then focused on the process.

I was initially wary of doing it programmatically since the slop risk was high. But in the end, thanks to the process, making doing it programmatically made me MORE confident. E.g. for many probs deterministic generators + exact/exhaustive validation was easier to audit than thousands of manually written ones (there are over 40k problems).

So I could leverage AI mostly after defining the rules and reviewing the system (which did have a lot of manual QA upfront especially for the different types, but went quickly after I built the QA tool). And that let me treat it more like a problem I'm better at solving as a product/systems guy, rather than a math expert.

TL;DR the problems were programmatically generated, but NOT LLM-generated/trusted. Rough process in case it helps you:

1. Define the problem families explicitly. Each has bounded inputs, known mathematical rule, answer type, constraints & presentation rules.

2. A deterministic compiler generated the problems. So, given the same source definitions/versions it'd produce the same corpus. So there's no AI inventing random questions at runtime.

3. Correct answers are computed from the underlying math, not from the rendered text. Integer/rational problems use exact arithmetic. More complex symbolic cases use validation recipes, incl. offline SymPy if needed.

4. LaTeX is just presentation (tried other way, didn't work). Internally the problem is represented as structured math, then laTeX/MathJax derives from that structure, so it does NOT generate a LaTeX string and then just hope it's interpreted correctly.

5. Generation/verification are separate steps. Every generated record has to pass mathematical, domain, schema, answer format & presentation checks before it can enter the corpus. Symbolic cases can be independently verified, for example by differentiating a proposed antiderivative.

6. Whole families are tested, not just samples. Finite spaces can be exhaustively enumerated. Larger spaces get property, boundary, invariant, and regression tests. It also runs corpus-wide audits for malformed questions, duplicate IDs, invalid answers, broken rendering, unreachable answer forms, etc.

7. The shipped artifact is tied back to what was verified. Versions/content digests bind source definitions, generated problem, validation result & runtime representation together. If something changes, it has to be revalidated rather than reusing old results.

One thing I'm especially loving about AI is that it lets you convert increasingly more problems to something you're good at, which makes it easier to approach/solve in novel ways. It let me do this with Mathy, which would have been untenable before. And I've been doing the same with robotics, which is exciting as a software/product guy.

gmays··on Show HN: Make math automatic with Mathy
You know, I hadn't actually considered iCould sync, thank you! I didn't like the idea of not syncing devices either, but couldn't justify just eating the storage costs since it's a free app.

But using iCloud would let users cover their own storage costs for syncing practice/progress/medal/challenge data.

I'll take a look at implementing this when I get some free time so I can have high confidence in the approach. Thanks again!

gmays··on Show HN: Make math automatic with Mathy
Thanks. Yeah, related but problem families are more separate from the knowledge graph.

To be clear, the knowledge graph here is different than the Math Academy sense. What they have is far more impressive since it's for learning and convers prerequisites, related topics, etc. Where mine is more specifically focused on automaticity, with the assumption you've already learned it and just want to maintain/improve recall.

The Mathy knowledge graph is roughly:

Domain > Topic > Concept > recall target > prompt variants

A recall target is basically the smallest piece of knowledge that gets its own spaced repetition state. So, for example, `7 × 8` can be one target. `8 × 7` is just another presentation of that same target, while something like `56 ÷ 7` is a separate inverse retrieval target.

A "problem family" is more of an authoring/generation construct layered onto that. E.g. is defines a bounded class of problems w/exact operand ranges, mathematical rules, expected answer forms, exclusions, presentation rules, etc. It can then deterministically produce valid problems for the relevant targets. So its not a 1:1 mapping between graph nodes <> problem families.

For the curriculum I tried to keep mathematical identity separate from curriculum placement. The internal graph is organized around coherent mathematical concepts and independently meaningful retrieval skills. Then grade/course views are mappings over that graph. It should roughly follow common core, but with some gaps since I only wanted to cover stuff doable in your head. I also covered add'l memorization topics to supplement the Math Academy courses I plan to do since I'll need those myself. I didn't get them all, but I know the MA team plans to add automaticity stuff, so I assume by the time I get to those they may already cover it anyway.

To clarify SymPy does not ship in the iOS app. It;s only used offline during the content build/validation process for the classes of symbolic math where it's useful. The app only ships the validated content. Runtime grading is bounded + local. E.g. there's no Python, SymPy, runtime AI, or unrestricted CAS running in the app.

I had to iterate a lot to get it performant, working 100% offline and at a manageable size with so much content. There were some compromises but it works reasonably well so far.

gmays··on Show HN: Make math automatic with Mathy
Thanks for letting me know, will get this fixed tonight.
gmays··on Show HN: Make math automatic with Mathy
Thank you! ANd thanks for sharing, that was a good overview and I had a similar experience.

To answer your question, I'm a product guy, so it was easier for me to start with the user app/mobile UX I wanted + constraints rather than starting with the cards.

Then I created a problem view in the dev version of the app that let me see every type of problem to see how it rendered, how input worked, etc.

Then with that plus the system around it (details below) gave me higher confidence in generating the 40,000+ problems across all the topics. It still has room for improvement, but happy with how it came out so far.

So, for the problems, they were programmatically generated, but NOT LLM-generated/trusted, so the process would give me the confidence:

1. Define the problem families explicitly. Each has bounded inupts, known mathematical rule, answer type, constraints & presentation rules.

2. A deterministic compiler generated the problems. So, given the same source definitions/versions itd produces the same corpus. so there's no AI inventing random questions at runtime .

3. Correct answers are computed from the underlying math, not from the rendered text. Integer/rational problems use exact arithmetic. More complex symbolic cases use validation recipes, incl. offline SymPy if needed.

4. LaTeX is just presentation (tried other way, didn't work). Internally the problem is represented as structured math, then laTeX/MathJax derives from that structure, so it doesnt generate a LaTeX string and cross fingers its interpreted correctly.

5. Generation/verification are separate steps. Every generated record has to pass mathematical, domain, schema, answer format & presentation checks before it can enter the corpus. Symbolic cases can be independently verified, for example by differentiating a proposed antiderivative.

6. Whole families are tested, not just samples. Finite spaces can be exhaustively enumerated. Larger spaces get property, boundary, invariant, and regression tests. It also runs corpus-wide audits for malformed questions, duplicate IDs, invalid answers, broken rendering, unreachable answer forms, etc

7. The shipped artifact is tied back to what was verified. Versions/content digests bind source definitions, generated problem, validation result & runtime representation together. If something changes, it has to be revalidated rather than reusing old results.

So, starting out I assumed doing it programmatically would be liability. But rather confidence comes BECAUSE it's programmatic. For a large class of problems deterministic generators + exact/exhaustive validation was easier to audit than tens of thousands of hand written questions. So could leverage AI for all that, just had to define the rules/review the system.

I'm very happy with the outcome, but in terms of process it far exceeded my expectations in what I learned about approaching problems like this.

I'm a very heavy AI coding using (I burn hundreds of billions of tokens a year!) so this was a fun way to validate my approach and test some new ones to build high quality experiences, particularly on mobile which is more of a taste thing. The decade plus of product experience really came in handy in guiding the AI here, which took a lot of iteration on the experience and trying different things. It was a blast, and I was generally surprised at how fast it came together. Thank you again.

gmays··on Show HN: Make math automatic with Mathy
Thanks, that's good feedback. I will look at options, thank you.
gmays··on Show HN: Make math automatic with Mathy
Thank you! Yeah, Justin Skycak's writing really turned me on to the benefits of automaticity. Once you're aware of it you start to notice everywhere you lack it and how much it'd help in raising your ceiling.
gmays··on Show HN: Make math automatic with Mathy
Thank you!
gmays··on Show HN: Make math automatic with Mathy
Thank you, and great writeup! I'm in a similar position and wrote about my journey a couple of years ago here: https://gmays.com/how-im-relearning-math-as-an-adult/
gmays··on Show HN: Make math automatic with Mathy
Thanks! Yeah, that was a dealbreaker for me since I didn't want to bother with AI slop. So it took a while and a ton of testing/validation, but happy with how it came out.
gmays··on Show HN: Make math automatic with Mathy
Thank you, I will look at that bug and get something deployed tonight! I only tested on Mac/iPhone, so I appreciate the feedback.
gmays··on AI Adoption Across the United States
The top AI use county in CA is Yolo county...
gmays··on Matt Mullenweg Overrules Core Committers; Puts Akismet on WP 7's Connector List
Agree. And the meta point, after reading through to the core committer channel on the WP slack is that it's clear he's now more involved in the project again and making decisions. I haven't been involved for years, but while I was it seems he had other priorities (understandable).

But the rapid changes from AI are an existential threat to the long-term viability of WP. Rather than bike shedding about something relatively trivial, they need to focus on the bigger issues, which it's apparent he's trying to do.

Interestingly, the culture that sustained WP over the last 2 decades may now be working against it. Culture is really hard to change, but he now seems to have his 'wartime CEO' hat on trying to do it, which is the right move.

gmays··on Moving from WordPress to Jekyll (and static site generators in general)
I didn't want to hassle with migrating my WordPress blog, so now just deploy it to Github > Cloudflare Pages so it's served statically (fast + secure). It's free too, wrote a blog post on it a couple years back: https://gmays.com/how-to-host-wordpress-sites-free/

But these days any new site I build is on NextJS since coding agents make it a breeze.

gmays··on Show HN: Lightwave – Real-time notes app, 3.5 years of hand-rolled JavaScript
Slick UI and well thought through, like the simplicity of the approach.

On the collab side, any limitations on simultaneous users? Like just a couple at a time or can handle a team?

gmays··on AI should write 50%+ of your code
Good point, it's a mix. The "it'll only get harder" is also because things are moving so fast and it takes time to learn (especially across teams) and change habits. No past paradigm has moved this quickly, which makes it hard to grok.

I also fully agree with "don’t overdo your investment into this generation of tools". IMO there are too many "cutting edge" tools trying to do all of this sexy stuff that'll be irrelevant in the next few months.

It's best to keep things simple with tooling. I push the edge on my general approach (99% of everything is AI coded) but conservative with my tools (pretty much only using Cursor now) to have at least some layer of stability. Otherwise stacking too many cutting edge things just feels too fragile, and will decay as AI improves, causing other issues. And this stuff is moving so fast and these companies are sufficiently motivated that the best things will make it into the tools, like plan/debug modes in Cursor.

I also feel that agentic coding is fast enough for now, so I don't even bother with multi-agent workflows. I still get a ton done and it's already at the edge of my ability to design coherently. Sure I could get 10X more code written in parallel with 10X more agents, but I can't design that fast, so it's just hurry up and wait with worse quality. And if that much code is needed I'm probably doing something wrong anyway.

gmays··on Tinkering is a way to acquire good taste
Same. This is a surprisingly simple recipe for a happy life and helps prevent lifestyle inflation. It reminds me of PG's "Keep your identify small" (https://paulgraham.com/identity.html).
gmays··on Apple and Amazon will miss AI like Intel missed mobile
That's fair, but it wasn't the point of the article because it's messy. Many would argue that core LLMs are 'trending' toward commodity, and I'd agree.

But it's complicated because commodities don't carry brand weight, yet there's obviously a brand power law. I (like most other people) use ChatGPT. But for coding I use Claude and a bit of Gemini, etc. depending on the problem. If they were complete commodities, it wouldn't matter much what I used.

A part of the issue here is that while LLMs may be trending toward commodity, "AI" isn't. As more people use AI, they get locked into their habits, memory (customization), ecosystem, etc. And as AI improves if everything I do has less and less to do with the hardware and I care more about everything else, then the hardware (e.g. iPhone) becomes the commodity.

Similar with AWS if data/workflow/memory/lock-in becomes the moat I'll want everything where the rest of my infra is.

gmays··on Apple and Amazon will miss AI like Intel missed mobile
OP here, good points.

Your comment on Intel is correct, but it's also true that TSMC could invest billions into advanced fabs because Apple gave them a huge guaranteed demand base. Intel didn’t have the same economic flywheel since PCs/servers were flat or declinig.

That's a good clarification on Amazon, running on commodity hardware with competitive pricing != competing on price alone. It would have been better to clarify this difference when pointing out that they're trying the same commodity approach in AI.

gmays··on Apple and Amazon will miss AI like Intel missed mobile
True, but Apple is a consumer hardware company, which requires billions of users at their scale.

We may care about running LLMs locally, but 99% of consumers don't. They want the easiest/cheapest path, which will always be the cloud models. Spending ~$6k (what my M4 Max cost) every N years since models/HW keep improving to be able to run a somewhat decent model locally just isn't a consumer thing. Nonviable for a consumer hardware business at Apple's scale.

gmays··on Apple and Amazon will miss AI like Intel missed mobile
I'm somewhat bullish on Google as well, they have the opportunity if they can figure out the product (which they are bad at) and they have the edge in cloud with their models + TPUs.

But your comment about the phone could have been about horses, or the notepad or any other technology paradigm we were used to in the past. Maybe it'll take a decade for the 'perfect' AI form factor to emerge, but it's unlikely to remain unchanged.

gmays··on Apple and Amazon will miss AI like Intel missed mobile
Right, but remember Microsoft was 'working on' mobile also. The issue is that they're working on it the wrong way. Amazon is focused on price and treating it like a commodity. Apple trying to keep the iPhone at the center of everything. Thus neither are fully committing to the paradigm shift because they say it is, but not acting like it because their existing strategy/culture precludes them from doing so.
gmays··on Analyzing Modern Nvidia GPU Cores
The special sauce:

> "GPUs leverage hardware-compiler techniques where the compiler guides hardware during execution."

gmays··on X’s director of engineering, Haofei Wang, has left the company
For context in response to the questions about "Why build on X?"

I've be working to relearn math and there happens to be a large group of others also doing the same with Math Academy and sharing daily updates on X.

I found this inspiring (especially as the lessons got harder) so I tweeting my updates too. But I also wanted a way to independently track my progress across math and other areas to see progress over time, even if I changed tools or stopped tweeting.

So that's the reason I built app on X: So my tweets get logged in a GitHub-like habit graph to show progress over time. It just pulled my bio/profile from X (login with X) and tracks my habit tweets. It's super simple, but meets my needs perfectly. My habit page: https://xtreeks.com/gabemays

I understand the questions around the long-term stability of the API, but I'm optimistic.

gmays··on X’s director of engineering, Haofei Wang, has left the company
I've started using X a lot more in the last few months since I built an app that let's you track habits with a tweet called Xtreeks (yeah, I know..).

I enjoy the product, but wish they'd spend more time making the core elements of the product work. For example, aspects of the API just don't work as expected, like for some reason search and mention endpoints do not have support for long form posts (>280 characters) enough though X supports posts with thousands of characters. The result is the API appears to work for some posts and just silently fails for others.

In addition to the API issues, we've struggled with inexplicable labeling/suspension and shadow banning, even on the personal account I've had for over a decade (seemed to be triggered by using my VPN). I understand the desire to control spam, but it seems excessive. Or if you do it excessively, at least provide adequate tools/support to request review.

On my app's X account I paid for both API access ($200/mo) and the Verified Org status ($2,000) and had a hard time getting support that took days to reply, when it did reply at all. And when the person replied they had nothing to do with the account label process, so weren't able to help, which was quite frustrating. It was fine since this was a little side project, but if this was a business at scale and I was paying that much in addition to ad spend I'd be furious.

Anyway, I know nothing about the Head of Eng or what's at the root of these issues, but I'm a big fan of X and hope they're able to fix these things. It's such. valuable tool. I'm even fine if it's pay to play, but if someone is on the higher tiers of your paid plans the support should be available when they need it.

Page 1 of 11Next →