HNHacker News
TopNewBestAskShowJobs

isaacfung

241 karma · joined January 13, 2023

submissionscomments
isaacfung··on DeepSeek OCR
Some models use vector quantized variational autoencoders to discretize images into sequences of discrete symbols from a fixed codebook.

https://grok.com/share/bGVnYWN5LWNvcHk%3D_572b4955-6265-4210...

isaacfung··on Show HN: Dia, an open-weights TTS model for generating realistic dialogue
What AI tools have you used recently? Have you verified if they all use models trained on copyrighted material with permission?
isaacfung··on Hoodmaps: Publicly Annotated City Maps
Something similar happened to Google Map in Hong Kong recently. A dozen schools had their names changed by some pranksters. It's surprisingly easy.

https://www.scmp.com/news/hong-kong/society/article/3279201/...

isaacfung··on Tutorial on diffusion models for imaging and vision
Also this blog post https://yang-song.net/blog/2021/score/
isaacfung··on Tutorial on diffusion models for imaging and vision
You may find the huggingface course more approachable

https://huggingface.co/learn/diffusion-course/en/unit0/1

isaacfung··on Web scraping with GPT-4o: powerful but expensive
How do you keep table structure?
isaacfung··on Inductive or deductive? Rethinking the fundamental reasoning abilities of LLMs
It sees embeddings that is trained to encode semantic meanings.

The way we tokenize is just a design choice. Character level models(e.g. karpathy's nanoGPT) exist and are used for educational purpose. You can train it to count number of 'r' in a word.

https://x.com/karpathy/status/1816637781659254908?lang=en

isaacfung··on Inductive or deductive? Rethinking the fundamental reasoning abilities of LLMs
The text is converted to embeddings after tokenization. The neural networwk only sees vectors.

Imagine the original question is posed in English but it is translated to Chinese and then the LLM has to answer the original question based on the Chinese translation.

It's a flaw of the tokenization we choose. We can train an LLM using letters instead of tokens as the base units but that would be inefficient.

isaacfung··on Diffusion models are real-time game engines
The possibility seems far beyond gaming(given enough computation resources).

You can feed it with videos of usage of any software or real world footage recorded by a Go Pro mounted on your shoulder(with body motion measured by some sesnors though the action space would be much larger).

Such a "game engine" can potentially be used as a simulation gym environment to train RL agents.

isaacfung··on 'Model collapse'? An expert explains the rumours about an impending AI doom
When people say "LLMs are not capable of innovation", what exactly do they consider as innovation? If LLMs are not capable of innovation on their own, what if we augment them with means to interact with the environment so they can obtain new training data?

e.g. The minecraft bot Voyager can explore the game environment and extend its skill library(stored as a vector database), is that considered as innovation? There are also systems like leandojo/alphaproof that discover new proofs and use LLM in non-trivial ways(not just naively predict the next token in one shot). Reinforcement learning algorithms like AlphaGo/AlphaZero use self play and use monte carlo tree search to learn to outperform humans. You can similarly use LLMs to generate actions and estimate state values(check the language agent tree search paper).

Most people use LLMs by prompting them with some additional context(chat history and data retrived from database) but there is nothing that stops us from continuously improving a LLM(either by modifying its weight or augmenting it with external database) by asking it to evaluate the task outcome/error message and feeding it back to the LLM. We can also ask it to just keep on generating new tasks to experiment with the environment/internet to get new knowledge.

isaacfung··on How I won $2,750 using JavaScript, AI, and a can of WD-40
bro, it's not just about the money. He's promoting himself, building up a portfolio, learning new skills and entertaining himself.
isaacfung··on Betting on DSPy for Systems of LLMs
We have translators. Doesn't mean we can't replace them with a cheaper, more accessible tool. That's the whole point of automation.

Reasoning stuff is not useless. They provably(according to benchmarks) improve the performance of coding and math related tasks.

isaacfung··on Betting on DSPy for Systems of LLMs
This repo has some less trivial examples. https://github.com/ganarajpr/awesome-dspy

You can try STORM(also from Stanford) and see the prompts it generates automatically, it tries to expand on your topic and simulate the collaboration among several domain experts https://github.com/stanford-oval/storm

An example article I asked it to generate https://storm.genie.stanford.edu/article/how-the-number-of-o...

isaacfung··on What happened to BERT and T5?
Aside from being used alone, T5 is also used as the text encoder of some recent multimodal models.

https://stability.ai/news/stable-diffusion-3-research-paper

https://t5tts.github.io/

Related discussion

https://www.reddit.com/r/StableDiffusion/comments/1c0by2y/wh...

isaacfung··on Overcoming the limits of current LLMs
I don't get why some people seem to think the only way to use a LLM is for next token prediction or AGI has to be bult using LLM alone.

You want planning, you can do monte carlo tree search and use LLM to evaluate which node to explore next. You want verifiable reasoning, you can ask it to generate code(an approach used by recent AI olympiad winner and many previous papers).

What is even "planning", finding desirable/optimal solutions to some constrained satisfaction problems? Is the llm based minecraft bot voyager not doing some kind of planning?

LLMs have their limitations. Then augment them with external data sources, code interpreters, give it ways to interact with real world/simulation environment.

isaacfung··on Architectural cross-section of Kowloon Walled City
There's a recently released kungfu movie whose story happened in the Kowloon walled city. The government is planning to relocate the movie set to the original address for exhibition.

https://www.youtube.com/watch?v=5uZmmTut7Ak

isaacfung··on The Illustrated Transformer (2018)
Maybe it's easier to understand in the format of annotated code

https://nlp.seas.harvard.edu/2018/04/03/attention.html

isaacfung··on Show HN: Voice bots with 500ms response times
There are way more text training data than voice data. It also allows you to use all the benchmarks and tool integrations that have already been developed for LLMs.
isaacfung··on Show HN: I made a puzzle game that gently introduces my favorite math mysteries
How do I know if it is the original "Where's Waldo" under the paper?
isaacfung··on Why we no longer use LangChain for building our AI agents
Some "agents" like the minecraft bot Voyager(https://github.com/MineDojo/Voyager) have a control loop, they are given a high level task and then they use LLM to decide what actions to take, then evaluate the result and iterate. In some LLM frameworks, a chain/pipeline just uses LLM to process input data(classification, named entitiy extraction, summary, etc).
isaacfung··on Why we no longer use LangChain for building our AI agents
I am not sure what you mean by "turn these one-shot APIs into Markov chains." To me, langchain was mostly marketed as a framework that makes RAG easy by providing integration with all kinds of data sources(vector db, pdf, sql db, web search, etc). Also older models(including initial chatgpt) had limited context lengths. Langchain helped you to manage the conversation memory by splitting it up and storing the pieces in a vector db. Another thing langchain did was implementing the react framework(which you can implement with a few lines of code) to help you answer multi hop problems.
isaacfung··on Transformers Can Do Arithmetic with the Right Embeddings
The current gen llms tokenize numbers digit by digit unlike earlier llms.
isaacfung··on Llama3 implemented from scratch
I recommend reading https://github.com/bkitano/llama-from-scratch over the article op linked.

It actually teaches you how to build llama iteratively, test, debug and interpret the training loss rather than just desribing the code.

isaacfung··on Ilya Sutskever: “If you learn all of these, you’ll know 90% of what matters”
Write a chrome extension that prompts a LLM/uses a custom trained (BERT based) text classification model to filter the posts.

https://github.com/thomasj02/AiFilter

https://github.com/GoogleChrome/chrome-extensions-samples

https://huggingface.co/docs/transformers/en/tasks/sequence_c...

isaacfung··on Are Japanese anime robots isometric or allometric?
There is an upcoming Mech movie Atlas, starring Jennifer Lopez with a heavy Titanfall vibe.
isaacfung··on Elon Musk sues Sam Altman, Greg Brockman, and OpenAI [pdf]
Open source would also mean it is available to sanctioned countries like china.
isaacfung··on Show HN: Supermaven, the first code completion tool with 300k token context
These may be useful

https://github.com/ggerganov/llama.cpp/blob/master/grammars

https://github.com/guidance-ai/guidance?tab=readme-ov-file#c...

https://github.com/eth-sri/lmql/issues/172

isaacfung··on LongRoPE: Extending LLM Context Window Beyond 2M Tokens
It may provide a way for AI to learn a complete new topic(a new programming language, stock investment, sport, games) in a zero shot way and immediately apply it by just feeding a youtube playlist/udemy course into an AI model.
isaacfung··on BASE TTS: The largest text-to-speech model to-date
There are lots of video content with audio. We can train a facial expression classification model to detect the speaker's emotion(we can also use a multimodal model to take in consideration of the language context).

Another potential source of data is voice acting script of animations. I always thought the storyboards of films/animations can be great annotated training data but it seems there are no open datasets, probably because of copyright issues.

isaacfung··on Antagonistic AI
It seems we can add antagonistic AI to correct many kinds of undesirable human behaviors

Browser extension that monitors how much time I have wasted on social media/porn

Browser extension that reminds me when I made a comment without reading a linked article(like right now)

Imagine a social media platform with an AI that will spot the most stupid comment and shame it using logical arguments and factual evidences, it can even dig up the user comment history to expose their hypocrisy

Page 1 of 4Next →