Jaithon 3, a fast programming language with the perfect syntax
github.com
github.com
Jaithon has Python features such as comprehensions, f-strings and first-class functions with declared fields, explicit visibility, traits and checked type annotations, along with syntax choices from lua, Java, bash, c++, go, and rust.
The compiler separates lexing, parsing, type checking and bytecode generation. The VM has 107 opcodes along with a JIT compiler to speed stuff up, polymorphic inline caches and a garbage collector.
Jaithon is nearly completely bootstrapped, with the lexer, parser, and bytecode generation built completely within Jaithon itself. The syntax of jaithon code is also easily customizable.
You can build and run it with:
git clone https://github.com/abhiramasonny/jaithon cd jaithon make ./jaithon examples/hello.jai
It would mean a lot if you star the project on my GH as I am trying to reach 15 stars soon :) anyways, lmk if you have any feedback. Currently Jaithon is between Java and C++ for speed (a more detailed benchmark exists within the project by running make benchmark) and I am in the process of optimizing the VM.
Would love to hear yalls thoughts!
- Abhi (abhiramasonny.com)
I have no reason whatsoever to use this, but it's a cool project!
error[E0301]: cannot assign to immutable binding `x` --> examples/demo.jai:7:5 | 5 | let x = 1 | - `x` declared immutable here ... 7 | x = 2 | ^^^^^ assignment to immutable binding | help: change the declaration to `var x = 1`
^^ That was shown from the Readme, more extensive examples are in documentation!
Is this what you were refering to or something else, I completely agree that erorr handeling is very important.
[0]: https://github.com/abhiramasonny/jaithon/blob/main/LANGUAGE....
Exceptions provide flexibility in error handling, and most code just forwards and you end up in the same situation anyway. Something like this (go style):
if result, err := do_thing(); err != nil { log.errorf("Got an error! %v", err); return err; }
Or the rust equivalent isn't super useful, the code is just outputting an inferior form of stack tracing and debugging.
Honestly unless the code at the immediate site of the error can handle the issue, or perhaps one level higher, having precise error information (usually obscured by some generic Error class anyway) isn't very helpful, and ends up propagating up to a high level where the whole thing is terminated / cleaned up anyway, which is what exceptions provide automatically.
Also, only catching the things that you can deal with and know about is useful in the sense it keeps code flexible (e.g. you don't handle a DiskFull exception because the only thing you can do with it is throw anyway, and if you had a DiskFull error returned, you would just return it up the stack / panic). As new unhandled errors emerge, you just throw them up the stack, the same way error code would, except errors just require you to explicitly manage the machinery everywhere, requiring rigid, over-specified, fragile code in many cases.
I do see the argument for explicit errors especially in system programming, realtime / perf-critical, kernels, etc. But this language doesn't appear to be targeted at that, and uses GC. So having exceptions seems like a valid design choice to me. Using them also frees you a bit since you can pass around functions, captured references, threads, etc. in a bytecode + GC lang without worrying to much about the error states and memory ownership.
At the point where you add a license, add an AI usage section to your readme. This is a perfect example.
The syntax is just a combination of my personal preferences :) seems like we share a bunch in common!
I see a value in defaulting to immutable unless explicitly stated, like in Rust. But having to choose between two different 3-letter words, like let and var, what stops one from always using var?
Coming to Jaithon in particular, I also wonder what the scope of immutability is, when it allows the following:
let names:list[str]=[]
names.push("whatever")
If a function takes a list[str] as input, how do I signal that it does or does not modify the list? And if it does, will it only shuffle the elements (e.g. sort) or will it also change the strings themselves (e.g. uppercase all elements)?Const-ness is a complex topic, with more cases than can be captured by two keywords at variable definition.
There is def nothing stopping you from using var everywhere, but (let) lets the compiler like catch any mistakes you did not mean to do. In jaithon, let is whats known as shallow: it prevents (in ur example) names = ... but the refrenced list is still mutable. Currently jaithon doesnt have anything like a read-only parameter so just knowing list[str] doesnt tell u anything about if a function mutates the list.
Good job tho! amazing prompting.
Anyways, back to your question, the architecture of jaithon is probably the most complex thing, and AI (atleast now) cant do this reliably. Syntax as well is something that is very abstract that AI isnt good at choosing. there is a lot more behind the scenes than simply oneshoting a prompt with opus and expecting it to work (spoiler it wont work well at all.)
Would love to see you try to do this with pure AI
"Most people on HN", yes, probably.
Even though people that don't work in tech can now build a lot of stuff I think most wouldn't even know where to begin with stuff like this.
Heck, I work in a web development/design agency with people who code in JS/TS all day and they still couldn't tell you the difference between interpreted and compiled code. Why? Because it's simply not relevant to their daily work.
"Make me a full implementation of a unique programming language." Then they QA, prompt, and continue the loop until they have something.
Understanding the code is NOT needed.
Programming languages are designed, not generated. What you're describing can get you an implementation for a thing, but getting to an implemented thing has never been the core issue in programming languages; indeed since the beginning of the field, implementations were nice things to have but the language itself was a different artifact. What was important were the properties of the language, which came from a careful consideration and balance of tradeoffs, and that's not something you get with "implementation of a unique programming language" because that can get you anything.
Languages are only interesting to people if the language has particular guarantees which are important for particular use cases important to the people. And we know this because before LLMs we also had a situation where anyone could write a programming language (and they did), so even before LLMS we had 1000s of low-effort implemented languages no one knew about or used. LLMs have accelerated this but it hasn't changed any of the fundamentals about how PLs are valued: the implementation does not matter, it's the language design, community, and ecosystem.
Then I iterate on this in a QA loop. I don't need to understand. I just need to test the high level details. The input and outputs and make sure everything works. The implementation details I don't need to understand at all.
I can tell the LLM to focus on whatever use case I want adjust the language to fit whatever preference I need. Even if the inception of the idea is blurry the LLM can crystallize a starting point and the "design" of the language can arrive piecemeal in the test and iterate loop.
> I don't need to understand. I just need to test the high level details. The input and outputs and make sure everything works/
If you want it to be something maintainable, no you can't do that. If you try, you'll get parallel interfaces, half baked APIs, a bunch of special cases that are incompatible and brittle, and eventually it'll have to be refactored and good luck with that -- the agents have a special failure mode there that is quite fun (building a artifact verification cathedral and then spending all its time verifying that instead of working on code).
> Even if the inception of the idea is blurry the LLM can crystallize a starting point and the "design" of the language can arrive piecemeal in the test and iterate loop.
This is why languages are designed. This process won't produce anything but mush.
> I can tell the LLM
Okay but have you though? Have you actually tried building a language this way? How long did you maintain it for? Can we see it?
Anyway, even if what you're saying is 100% true, all that means is the bar is higher. Language implementation would mean nothing in that world, design would be everything. Which, that seems... fine? But it's not the current world IME.
I've tried it too. It worked.
>If you want it to be something maintainable, no you can't do that. If you try, you'll get parallel interfaces, half baked APIs, a bunch of special cases that are incompatible and brittle, and eventually it'll have to be refactored and good luck with that -- the agents have a special failure mode there that is quite fun (building a artifact verification cathedral and then spending all its time verifying that instead of working on code).
You're just saying this. You haven't actually tried having an LLM write all the code and have the LLM do the maintenance. Your statement lacks any real world data and it's just wishful thinking. The answer here is actually quite complex because we have people like you who complain about it and say they "know" it won't work because they've seen it with their own eyes while people like me have seen it totally work and be totally doable as well.
So who's reality is true? Only time will tell. But I find it unlikely you've tried to have an LLM try to maintain an active project in prod. I have.
>This is why languages are designed. This process won't produce anything but mush.
Typescript is the result of years of iteration on a tiny language called javascript.
>Okay but have you though? Have you actually tried building a language this way? How long did you maintain it for? Can we see it?
Nope. haven't done it. But I have done things more complex then language design. And I can't show you because my employer owns it. But FYI language implementation follows very common patterns and this makes it very very easy for LLMS to create one. A beta of a language can be done in about a week.
An experienced human can write code better than an LLM for low level tasks almost 100% of the time, the only problem is its REALLY slow for a singular person to write code when compared to like an LLM running at 600 tokens per second or something. The best combination of the two is (in my opinion) running claude on manual mode cus if it does something really bad you can either fix it yourself or reprompt it, but there is a LOT of active things you have to be doing.
>Nope. haven't done it. But I have done things more complex then language design. And I can't show you because my employer owns it. But FYI language implementation follows very common patterns and this makes it very very easy for LLMS to create one. A beta of a language can be done in about a week.
its not about the fact that the thing you are doing is "more complex" (in my humble opinion the only thing more complex than language design is building an OS) but its more of the fact that if the LLM writes one line of code that is not like optimal for performence, your language will like leak memory or it will run really slowly, or something will happen suboptimally. you cant just leave it autonomously for large periods of time (I have tried and regreted this) but you can still use it, as long as you read over and esnure that the code is optimal. if you dont do that we are just moving backwards in time and the code you write is very unmaintanable, slow, leaks memory or is volatile, and impossible to scale up or down.
Also am slightly dismayed that new langs are copying the f-string, which was chosen only because Python had no more Ascii punctuation chars left for such a feature. Would much recommend shell-style quote syntax instead, perhaps omitting $.
Syntax is a personal choice. Can always be changed. Architectural choices are difficult to manage later on.
The VM seems to be a stack VM. I have my own python replacement project (https://news.ycombinator.com/item?id=48090665) that I am developing privately for now. LLM-assisted ofc. Started with a bytecode VM and then switched over to a register one. The biggest break from regular language systems was the decision to move to an Erlang-style preemptive scheduler.
Also, I was curious why you swapped over to a register based VM, could you give more clarity on that and the benifits/negatives? currently im using some sort of bytecode vm, though its not like im married to the idea of a bytecode vm.
Also (dont mean this as any sort of advertisement or anything like that) would you mind starring the repo? trying to hit 15 stars :D dont feel forced or anything like that.
Not the OP but there are real performance benefits, I've been poking at a wasm VM and it has two jit backends where one is pure copy-and-patch while the other caches the locals in registers using the function args + copy-and-patch and there is a significant performance gain just from that alone. A push/pop from a stack is fairly expensive while the register caching keeps things in the CPU's happy place. The smallest gain was ~2x over the interpreter on memory bound tasks while the largest was ~20x on math heave kernels. Admittedly, the interpreter isn't the fastest thing ever as its one and only goal is conformance with the spec to use for differential testing but the difference between the the two jit levels are somewhere in the neighborhood of 1.5-5x depending what the code is up to.
The three biggest performance gains, from the random benchmarks, are quality of the bytecode out of the compiler, the jit itself and register caching from what I can tell from the fancy chart I had Claude make and a good squint. Tail-calling would be somewhere on that list too but I can't measure that as all the opcode do the tail-calls between each other as that's just how it was all put together, the code the interpreter runs is the same code the copy-and-patch jit stitches together as they are both generated from the same DSL. Which is also the biggest cost with the register caching as the code template file grew from tens of kilobytes for the 407(?) wasm opcodes to ~3MB for all the specialized ones to pass the locals in eight args but that's really just a binary size thing, the stitched together functions just pick and chose the ones they need.
Long winded way to say CPUs like when you keep things in registers, I suppose...
Jaithon wasn't necessarily based on a particular language or multiple languages even per say, its just my preferences from all the languages ive coded in the past :)
[1] https://github.com/orgs/community/discussions/196787#discuss...
Maybe one day I can plead with the github gods to accommodate for jaithon :)
How would the human ever know that, though?
> Additionally, around 80% of the raw code in this repository was produced with agentic coding tools
Not to be a party pooper, but even if the author wouldn't themselves describe it as "slop," the term still applies. It isn't really a "collaboration" if the vast majority of the implementation was carried out by one (non-human) party.
In my personal opinion (you are 100% free to disagree as this is just an opinion and not fact) the 'hard' part of programing is coming up with the proper architecture, and ensuring the codebase is maintainable and is able to be built on top of. Additionally, the word slop in my opinion signals that the product is terrible. In my opinion, the product of Jaithon is not terrible (like some AI generated code is) and as I have not only been highly transparent with the use of AI and also used it with an appropriate maner with limits, by my definitions of it, this is not considered slop.
Additionally, the vast majority of the CODE was carried out by AI, not the implementation. Architecture & language automota were actually created in 2022-2023 while I was in 8th grade. AI was only used rather recently to speedrun development.
Hope this serves as further clarification by what my readme and my words ment.
> It isn't really a "collaboration" if the vast majority of the implementation was carried out by one (non-human) party.
So then there is a ratio at which you would say it's a collaboration? What if I write the function signatures and allow the AI to fill in the implementations? What if I write half the implementation and the AI fills in the other half? What if I write comments and the AI fills in an implementation of the thing and then I review it and fix it manually, is that vibecoding or slopcoding? Are vibecoding and slopcoding the same?
Moreover: does any of that even matter if the end result is functionally the same either way?
I'm just trying to understand the acceptable way to use these tools because the opinions people have range from "not at all, never" to "all the time, always".
i think this is an appropriate use of AI but would love to hear other opinions
This line > Additionally, around 80% of the raw code in this repository was produced with agentic coding tools
came prior to this line > I would not consider myself a "vibecoder", or jaithon as "ai slop"
Phrasing it the other way like you just did makes it seem like I contradicted my self heavily, however in reality I had the second line with a bunch of reasoning in the readme.