There are many research papers on models using characters directly. One challenge is that effective context length is smaller.
1,147 karma · joined March 26, 2023
There are many research papers on models using characters directly. One challenge is that effective context length is smaller.
The model plans (with thoughts visible), can solve complex problems with Flash speeds, and more ..."
- Logan Kilpatrick
https://discord.com/login?redirect_to=%2Fchannels%2F11894982...
Aider: https://github.com/paul-gauthier/aider
It is state of the art on SWE-Bench and SWE-Bench Lite. https://aider.chat/2024/06/02/main-swe-bench.html
Earlier thread. https://news.ycombinator.com/item?id=39494760
For professional cards, I've noticed dihuni.com has good prices. I have never purchased from them and have no idea what dealing with them is like.
The Unreasonable Effectiveness of Eccentric Automatic Prompts
It is only a matter of time before you and your company are affected by the pending regulations. In the future, almost all software products will be using AI models, in the same way that most software products use the Internet today, whereas they did not in the 1990s.
Imagine you had to license Oracle software because MySQL or PostgresSQL could not offer certain capabilities, or are less capable because of regulation.
Now also imagine that your products have to agree with the political world view of either Sundar Pichai or Elon Musk. And if you need capabilities only present in one of the commercial alternatives, you don't even that that choice.
https://www.ntia.gov/federal-register-notice/2024/dual-use-f...
Fact Sheet: https://www.whitehouse.gov/briefing-room/statements-releases...
Full Details: https://www.whitehouse.gov/briefing-room/presidential-action...
Training is both memory throughput and compute constrained. Much research in speeding up training goes into optimizing HBM to SRAM communication. The equivalent for your chips would be communication from the SRAM of one chip to the SRAM of another, where it sounds like your architecture has a major memory throughput advantage over GPUs. So I assume you don't have a proportional compute advantage?
By the way, it's great to see a non von Neumann architecture showing a major performance advantage in a real world application. And your chips are conceptually equivalent to chiplets; you should have a major cost advantage on bleeding edge process nodes if you scale up manufacturing. Overall very impressive!
Overview https://hazyresearch.stanford.edu/blog/2023-12-11-zoology0-i...
Zoology 2 https://hazyresearch.stanford.edu/blog/2023-12-11-zoology2-b...
Monarchs and Butterflies: Towards Sub-Quadratic Scaling in Model Dimension. Scaling in model dimension as opposed to sequence dimension scaling in the previous posts. https://hazyresearch.stanford.edu/blog/2023-12-11-truly-subq...
HashGraph (2019) https://arxiv.org/abs/1907.02900
Anyone know what the most performant CUDA hash table implementations are these days?
https://www.bloomberg.com/news/articles/2023-11-18/openai-al...
Chang does not make it clear whether the source is close to Sam Altman.
A few weeks ago, I had spotty service with Bing Chat where it would keep resetting the conversation which I assumed was due to load. In general all these LLM services are in constant flux because they are tuning both the models and UI. They feel like alpha quality products in terms of stability.