HNHacker News
TopNewBestAskShowJobs

JustinAngel

19 karma · joined March 17, 2017

submissionscomments
JustinAngel··on [dead]
Last month I hosted an in-person workshop about building your own large language model without any math or ML prerequisites. It covers everything from machine learning fundamentals, deep neural networks, transformer architecture, and pre/post-training. I’m releasing recordings and training materials for you to watch!

>> https://go.justinangel.ai/video-1 <<

The workshop’s goal is to grok all parts of modern LLM development. Each section of the workshop has slides teaching the concepts, followed by excel-by-hand exercises developing intuition for the math, and then coding tutorials. One participant, Emily HK, noted: “The best part of this workshop is that all the content is still available online for me to refresh my memory at any point”. Well, now you have access to these materials as well!

* 23 Workshop Videos @ https://go.JustinAngel.ai/playlist

* 250-page Slide Deck @ https://go.JustinAngel.ai/deck

* 50 Excel and Code Exercises @ https://go.JustinAngel.ai/drive

YOUTUBE LINKS

1. Sampling Large Language Models https://go.justinangel.ai/video-1

2. Reverse Engineering Large Language Model https://go.justinangel.ai/video-2

3. Perceptrons: wx+b https://go.justinangel.ai/video-3

4. Activation Functions: ReLU, GELU, SwiGLU https://go.justinangel.ai/video-4

5. GPU Coding: PyTorch, torch.compile(), fused kernels, CUDA, Triton https://go.justinangel.ai/video-5

6. MLPs/FFNs: Multi-input, Multi-Layer Perceptrons, Feed-Forward Networks https://go.justinangel.ai/video-6

7. Loss Functions: Residual errors, RMSE, Cross Entropy, Loss Landscapes https://go.justinangel.ai/video-7

8. Backpropagation: Training loops, Optimizers, Learning Rate, Batch Size https://go.justinangel.ai/video-8

9. Saving & Loading Models https://go.justinangel.ai/video-9

10. Initialization: Kaiming, Glorot https://go.justinangel.ai/video-10

11. Residuals: Addition, Scaling, Gated, Concatenation https://go.justinangel.ai/video-11

12. Normalization: Pre-norm vs. Post-norm, RMSNorm, BatchNorm, LayerNorm https://go.justinangel.ai/video-12

13. Regularization: Dropout, Gradient Clipping, Weight Decay https://go.justinangel.ai/video-13

14. SoftMax https://go.justinangel.ai/video-14

15. Tokenizers: By Character, By Word, BPE, SentencePiece https://go.justinangel.ai/video-15

16. Embeddings: Absolute vs. Learned, Sinusoidal vs. RoPE https://go.justinangel.ai/video-16

17. Attention: MHA, GQA, MQA, MLA https://go.justinangel.ai/video-17

18. Transformers https://go.justinangel.ai/video-18

19. Pre-training: Data Sources, Datasets, HTML Cleaning, Quality Filtering, Sharding https://go.justinangel.ai/video-19

20. Evaluation: Leaderboards, Benchmarks, Verifiers vs LLM-as-Judge https://go.justinangel.ai/video-20

21. Instruction Tuning: Alpaca & Other Formats, Self Instruct, Capabilities https://go.justinangel.ai/video-21

22. Reinforcement Learning: Policy Optimization, SimPO https://go.justinangel.ai/video-22

23. What We Didn't Cover: Scaling https://go.justinangel.ai/video-23

JustinAngel··on [dead]
Hi internet friends, I recorded a workshop about building your own LLM without any math / ML prerequisites. It covers everything from machine learning fundamentals, deep neural networks, transformer architecture, and pre/post-training.

The only prerequisite is being comfortable with learning through code & excel examples.

1. Sampling Large Language Models https://go.justinangel.ai/video-1

2. Reverse Engineering Large Language Model https://go.justinangel.ai/video-2

3. Perceptrons: wx+b https://go.justinangel.ai/video-3

4. Activation Functions: ReLU, GELU, SwiGLU https://go.justinangel.ai/video-4

5. GPU Coding: PyTorch, torch.compile(), fused kernels, CUDA, Triton https://go.justinangel.ai/video-5

6. MLPs/FFNs: Multi-input, Multi-Layer Perceptrons, Feed-Forward Networks https://go.justinangel.ai/video-6

7. Loss Functions: Residual errors, RMSE, Cross Entropy, Loss Landscapes https://go.justinangel.ai/video-7

8. Backpropagation: Training loops, Optimizers, Learning Rate, Batch Size https://go.justinangel.ai/video-8

9. Saving & Loading Models https://go.justinangel.ai/video-9

10. Initialization: Kaiming, Glorot https://go.justinangel.ai/video-10

11. Residuals: Addition, Scaling, Gated, Concatenation https://go.justinangel.ai/video-11

12. Normalization: Pre-norm vs. Post-norm, RMSNorm, BatchNorm, LayerNorm https://go.justinangel.ai/video-12

13. Regularization: Dropout, Gradient Clipping, Weight Decay https://go.justinangel.ai/video-13

14. SoftMax https://go.justinangel.ai/video-14

15. Tokenizers: By Character, By Word, BPE, SentencePiece https://go.justinangel.ai/video-15

16. Embeddings: Absolute vs. Learned, Sinusoidal vs. RoPE https://go.justinangel.ai/video-16

17. Attention: MHA, GQA, MQA, MLA https://go.justinangel.ai/video-17

18. Transformers https://go.justinangel.ai/video-18

19. Pre-training: Data Sources, Datasets, HTML Cleaning, Quality Filtering, Sharding https://go.justinangel.ai/video-19

20. Evaluation: Leaderboards, Benchmarks, Verifiers vs LLM-as-Judge https://go.justinangel.ai/video-20

21. Instruction Tuning: Alpaca & Other Formats, Self Instruct, Capabilities https://go.justinangel.ai/video-21

22. Reinforcement Learning: Policy Optimization, SimPO https://go.justinangel.ai/video-22

23. What We Didn't Cover: Scaling https://go.justinangel.ai/video-23

Each section has slides teaching the concepts, followed by excel-by-hand developing intuition for the math, and then coding examples. The goal is able to grok all parts of modern LLM development.

We did this workshop in-person in San Francisco last month and hopefully the spaciousness of watching online works for everyone. https://emilyhk.com/llm-workshop/

If don't like watching videos, you can get the slides and exercises and work self-paced. https://go.justinangel.ai/deck

JustinAngel··on [dead]
Hi HN, thought I'd drop a link to my thesis on developing clinically-effective AI psychotherapy. I wrote this paper for anyone who's interested in creating a mental health LLM startup and developing AI therapy.

Summarizing a few of the conclusions in plain english:

1) LLM-driven AI Psychotherapy Tools (APTs) have already met the clinical efficacy bar of human psychotherapists. Two LLM-driven APT studies (Therabot, Limbic) from 2025 demonstrated clinical outcomes in depression & anxiety symptom reduction comparable to human therapists. Beyond just numbers, AI therapy is widespread and clients have attributed meaningful life changes to it. This represents a step-level improvement from the previous generation of rules-based APTs (Woebot, etc) likely due to the generative capabilities of LLMs. If you're interested in learning more about this, sections 1-3.1 cover this.

2) APTs' clinical outcomes can be further improved by mitigating current technical limitations. APTs have issues around LLM hallucinations, bias, sycophancy, inconsistencies, poor therapy skills, and exceeding scope of practice. It's likely that APTs achieve clinical parity with human therapists by leaning into advantages only APTs have (e.g. 24/7 availability, negligible costs, non-judgement, etc), and these compensate for the current limitations. There are also systemic risks around legal, safety, ethics and privacy that if left unattended could shutdown APT development. You can read more about the advantages APT have over human therapists in section 3.4, the current limitations in section 3.5, the systemic risks in section 3.6, and how these all balance out in section 3.3.

3) It's possible to teach LLMs to perform therapy using architecture choices. There's lots of research on architecture choices to teach LLMs to perform therapy: context engineering techniques, fine-tuning, multi-agent architecture, and ML models. Most people getting emotional support from LLMs like start with simple prompt engineering "I am sad" statement (zero-shot), but there's so much more possible in context engineering: n-shot with examples, meta-level prompts like "you are a CBT therapist", chain-of-thought prompt, pre/post-processing, RAG and more. If you're interested in reading more, section 4.1 covers prompt/context engineering, section 4.2 covers fine-tuning, section 4.3 multi-agent architecture, and section 4.4 ML models.

4) APTs can mitigate LLM technical limitations and are not fatally flawed. The issues around hallucinations, sycophancy, bias, and inconsistencies can all be examined based on how often they happen and can they be mitigated. When looked at through that lens, most issues are mitigable in practice below <5% occurrence. Sycophancy is the stand-out issue here as it lacks great mitigations. Surprisingly, the techniques mentioned above to teach LLM therapy can also be used to mitigate these issues. Section 5 covers the evaluations of how common issues are, and how to mitigate those.

5) Next-generation APTs will likely use multi-modal video & audio LLMs to emotionally attune to clients. Online video therapy is equivalent to in-person therapy in terms of outcomes. If LLMs both interpret and send non-verbal cues over audio & video, it's likely they'll have similar results. The state of the art in terms of generating emotionally-vibrant speech and interpreting clients body and facial cues are ready for adoption by APTs today. Section 6 covers the state of the world on emotionally attuned embodied avatars and voice.

Overall, given the extreme lack of therapists worldwide, there's an ethical imperative to develop APTs and reduce mental health disorders while improving quality-of-life.

JustinAngel··on [dead]
Repost since last post got "[flagged]".
JustinAngel··on How I lost 100lbs in 6 months
Yep. I called out my favourite one (warermelon) on that list and also called out the Weight Watchers list of "Power Foods" that's mostly comprised of fresh fruits and vegetables.

My issues with a blanket recommendation of fresh fruits and vegetables are around (1) perception as calorie-neutral, (2) prep time and (3) overall caloric variance.

1. Regarding perception, for me it's very easy imagining apples are "free" (= 0 calories) so I can just eat as many as I'd like. That's not true because an apple is approximately 70 calories. So in a universe where I think apples are calorie neutral (e.g. Weight Watchers has 0 PointsPlus apples) I might be overindulging in apples.

2. Prep time. Most vegetables require either a certain presentation or preparation. That time investment is a "barrier-to-eating" and could cause me to consume higher calorie foods and that's why I prefer very low prep time snacks. Nature's packaging requires some assembly in most cases.

3. The caloric variance amongst is meaningful. e.g. Celery being at 16 calories/100g and potatoes being at 77 calories/100g. That caloric variance there cautions me from saying carte-blanche "fresh fruits and vegetables work for my kind of diet".

JustinAngel··on How I lost 100lbs in 6 months
> "At this point I should say that I’m not a healthcare professional and I’m not recommending my dietary changes to anyone. I was in intense emotional pain every single day and made the decision to capitalise on it."
JustinAngel··on How I lost 100lbs in 6 months
Thanks for asking!

> "It was fairly easy due to a combination of low-grade depression that reduced my appetite and having fairly immense energy stores in the form of fat."

For me, it's all about focusing on the sensation of being full. I'd like to avoid the mental sensation of hunger as it'll be counterproductive to long term weight loss. Whenever I'm hungry I eat something substantial, for example: a quarter of a watermelon (~1 lb, 200 calories) or two bags of Miracle Noodles (~1 lb, 0 calories) with yakiniku/pasta sauce (100-200 calories) or a double bowl of miso soup (100 calories).