149 karma · joined June 26, 2017
Can you order any of these devices online as a regular person? Anybody can order a $300 Nvidia GPU and program it. This is the reason why deep learning originated on the GPUs. Forget those other AI accelerators, even if you bought something like a consumer grade AMD GPU, you couldn't program it because it's restricted. The reason why Nvidia's competitors are struggling is because their hardware is either too expensive or hard to buy.
Right now, I am grappling with RSI and I haven't been able to program for more than a few days in the past month. I am not even typing this, but using the Voice Access feature of Windows 11 in order to input this. I ordered an ergonomic keyboard (Glove80) and I am waiting for it. I also have an ergonomic mouse and even got an ergonomic chair. But I know that regardless of the case, I won't be able to program. for at least a few months. Until my hand recovers. If it does at all.
I am definitely going to check out this extension. Quick question: does it work on any language or just something like Javascript?
If they can get the manufacturing process for them worked out. I really do wish memristors turned out to be a success, HP hyped them so massively in the 2011-2015 era. What happened was that the material they had was susceptible to rusting, so what seemed like a good initial yield would become unusable some months down the road.
I google for memristors sometimes, and all the activity regarding them is still confined to the lab unfortunately.
I'd really like an avenue to get into the US market as a remote worker, but am being unfairly treated by this job market. It is a pity as I am both a highly skilled programmer and have nearly a decade of experience. I'd consider this service if it could serve to showcase my skills, but if I am not going to get any credit for doing the work personally, there doesn't seem to be much point to it.
Here is the full text.
Regexps aren't even Turing complete as far as I know, if whatever they have in their paper works for arbitrary programs it would be shocking. I'll give it a read.
*Edit*: The algorithm in the paper is a DP like algorithm for building regexes. They use a matrix, and it has all the potential strings to be checked on one axis, and all the potential regex programs on the other axis, and in-between values (the actual matrix values) are booleans saying whether the string matches the program. The algorithm builds the matrix iteratively.
I haven't understood how regex evaluation is done, probably directly, but obviously this algorithm is only for checking whether a particular regex program matches an output rather than general purpose synthesis.
We'll have to wait for AI chips to really scale genetic programming, GPUs won't cut it.
Remote: Only
Willing to relocate: No
Technologies: .NET (F#) mainly, but also Typescript and Python
Résumé/CV: https://docs.google.com/document/d/e/2PACX-1vTIQlCysV9y-w-PE...
My story is that even though I've always been talented at programming, to the point of winning the national programming championship of Croatia back in 2002, I didn't start it seriously until 2015. I had a dream of wanting to pursue the Singularity, and I worked hard every day, weekends included, to get closer to it. I wanted to get better at ML, so I worked on ML libraries in F#, which eventually grew into work on my own programming language Spiral. It took years of full time work, but I implemented my own PL, and a GPU based ML library from scratch. You can check out the language on the VS Code marketplace. This experience made me really good at implementing ML papers and algorithms. It also made a master of functional programming, and a strong generalist. Most of what I know isn't on the resume, and I am even familiar with dependently typed programming in languages like Agda, and theorem provers like Coq.
Unfortunately, I've stumbled. What I really wanted to do using these skills is create something like a poker bot, and crush the online gambling dens, but no matter how much effort I put into ML, I could never overcome the state of the art in a significant way. This is a huge problem since I am interested in RL, but RL only works on toy problems. I hate that I know everything that is wrong with current ML techniques, but have no idea how to fix it.
Instead of the world we have in the current timeline, it was obvious to me from the start that the way ML is currently done is broken, but I thought the research community would be able to overcome it, and give me some tools I could use to make a real world effect. Also, I thought there would be a ton of AI chip startups coming out with novel hardware, in which Spiral could have found its niche, but so far they've been a huge dud, and NVidia reigns supreme.
I am really looking for work so I can finance my own ML research, thus far I didn't need it, but I've come to the conclusion, that if I want to make a real breakthrough, I should be implementing genetic programming systems on AI hardware (not GPUs) but that approach would be highly intensive on computational resources which I cannot afford. In my work on Spiral, I've pretty much reinvented the field of partial evaluation, and if the job involved making advanced software like interpreters on hardware that had poor software support, I'd probably be peerless at that, even compared to anybody else in the world.
But as for web dev jobs which I am seeking currently, it feels like I am mid range in terms of skill. In the past few months I've gotten familiar with React, Fable, and now am pivoting to Blazor, and by the end of the year, I should be good at that.
Right now I am making Youtube videos of using F# for webdev that you can check out here: https://www.youtube.com/channel/UC6e__mOHSUPTu9HQ2W2RGcQ
I am also a prolific (if not a popular) writer: https://www.royalroad.com/fiction/57747/simulacrum-heavens-k...
My English proficiency is titled way too hard towards the writing side, so I've been doing screencasting in part to surmount that weakness.
Also seconding that other post. Actor models and async concurrency are only useful if you need to send messages between machines, but otherwise you want to use synchronous concurrency as it is easier to deal with.
Better algorithms would require writing new ML libraries which would weaken the stranglehold GPUs have on ML. New niches opening would inevitably be bad for Nvidia, but good for upstarts.
The backend I am demoing here is for UPMEM, but now that I have it as a template, I could easily make similar backends for other kinds of devices assuming I have access to at least a simulator.
Let me plug my own, though it is not a short story. It is about a kid who is LARPing as an unaligned AI.
Not related to the Aluminum Scandium Nitride memristors from the current paper, but this article I just found fascinating. I've been a stealth fan of memristors for a long time, and this finally answers why HP's efforts failed in such a crazy fashion. It is because they were using metal oxide memristors, and they rust easily.
Maybe if the devices from the current paper turn out to be good that could finally act as the fuel for the next AI wave after deep learning? The paper looks good to my eyes, but I am not a materials expert.
Regular function (not heap allocated closures) do not have a runtime footprint. Their body just gets tracked by the partial evaluator, but never manifests unless they are applied. Their variables (usually stack-alloc'd) get tracked on an individual basis.
So it is not the case that functions are stack allocated. They are fully known.
A closure is something that is opaque to the partial evaluator. Closures do need to be heap allocated because they could be recursive data types.
> "expanding a function body into its call sites to elide function-call overhead"
If you do this, not only do you not need the function call overhead, but you also do not need the overhead from heap allocating a function a runtime (converting it into a closure.)
I am not sure how good I am at explaining this. I would appreciate somebody going over the docs and giving me some feedback on it. When the last Spiral was posted on HN, there was some initial excitement for a few days and then it was crickets.
If you are willing to give it a try, I'd be happy to answer all your questions in the Spiral issues.
I said it is an equivalence in terms of compilation because, whether it is allocated as a closure on the heap, or tracked at compile time has no bearing on the correctness of the program. It only affect its memory allocation and performance profile.
If you are a C programmer, or working at a similar level of abstraction, this concern over heap allocations might seem academic, but to a functional programmer such as myself there is a lot of importance because we use function composition for all of our abstractions and don't want them to heap allocate as closures. The more abstract the code we are writing is, the more inlining matters.
The way I am describing this is confusing because inlining is not an event that happens. Rather inlining is the default. Tracking functions at compile time is the default. Heap allocation of them and their conversion into closures is an event, after which the compiler stops tracking variables of a function on an individual basis and keeps track of them at a lower resolution.
> I would have thought the opposite of heap allocation was "stack allocation"?
| Stack | Heap
------|-------------
Full | x |
Box | | xTo get a full picture of how Spiral's partial evaluator sees functions and recursive unions, it consider the table above. Spiral is designed so that it is very easy to go between the fully known and boxed. You don't actually control heap and stack allocations explicitly, but instead shift the perspective of the partial evaluator through the dyn patterns and control flow.
It just so happens that a very natural way of generating code for fully known functions is to leave their variables as they are on the stack. If the function needs to be tracked at runtime, then those variables are used to make a closure. The natural way of compiling things that are fully known (functions, unboxed unions) is to put them the stack. While things that are boxed (closures, boxed recursive unions) are on the heap.
This is just on the F# backend. If I were compiling to LLVM for example, I would not use a stack, but the infinite amount of virtual registers instead. And non-recursive union types are compiled as structs, meaning they should be on the stack even in boxed mode.
This view of memory is more complicated than the usual heap vs stack allocation one, because I am playing a game here where I attach a bunch of other concepts to it, but it is very useful and simplifies actual programming.
The docs on the main page are what you want.
It has things like structural recursion (similar to Agda), dependent pattern matching (the biggest benefit of which would be proper variable naming), unicode, `calc` blocks, good IDE experience (it actually has autocomplete) with VS Code (I prefer it over Emacs and the inbuilt CoqIDE is broken on Windows), mutually recursive definitions and types, and various other things that are not at the top of my head.
If I were to sum it up, the biggest issue with Coq is that it does not allow you to structure your code properly. This is kind of a big thing for me as a programmer.
https://thehumanevolutionblog.com/2015/01/12/the-poor-design...
Why was this the case?
I definitely agree with them in this matter. Macros might have a role in language development, but they should not be a stand in for compiler optimizations. For safety and speed, the type system should be there.
https://github.com/mrakgr/The-Spiral-Language/blob/0741959ac...
Here is how the example you've shown could be done in Spiral. Maybe I should add express support for literal testing in pattern matching, but it is not a pattern that comes up too often by itself.
> I feel like what Julia does is just more low level right now - And you get pretty far with just multiple dispatch!
What you say is exactly right as pattern matching compiles down to those low level operations, so there no reason at all why those low level operations should be done by hand. Pattern matching does not depend on a particular type system or whether the language is dynamic or static. There are only so many good ideas in programming languages and this is one of them.
Though there is some overlap between pattern matching and multiple dispatch, the roles are different. The purpose of multiple dispatch is extensibility, but the purpose of pattern matching is destructuring.
https://stackoverflow.com/questions/2502354/what-is-pattern-...
Note the great disparity in the F# examples that do pattern matching and the C# examples that do manual reflection.
https://github.com/mrakgr/The-Spiral-Language/blob/0741959ac...
Here are a few very simple examples of it in action in Spiral. I use them as compiler tests.
https://github.com/mrakgr/The-Spiral-Language/blob/0741959ac...
Here is quite a complex example of how it is used in action. I won't go into detail of this here, but you can see how I repeatedly match on the contents of a module at different times in order to get more generic functionality for the kernel.
The particular kernel shown here is the most complex one that exists in the library right now - I am yet to get to things like generic matrix multiplication and convolution, but I'll get there eventually.
> I see how it sounds nice, but I can't really imagine right now how it would improve my life.
Let me just say that it is really difficult to know ahead of time how a particular language feature would affect your programming life. I could have said the same thing about first class functions back in 2015. I am sure in the future there will be such features I can't even imagine right now.
> And if we don't get it as part of the compiler, we can definitely do things like that with https://github.com/jrevels/Cassette.jl/
I'll have to watch the talk on this. Thanks for the link.
Then why do I nowhere see that actually being done? Reflection on tuples should be done much like pattern matching on lists in functional languages. What is the point of that VarArgs nonsense? I know that Julia has pattern matching - now it needs to take the next step and actually make use of it.
Is this Julia's way of making itself familiar to C programmers?
I see those `...` elipses used as some kind of operator in both C++ (and Racket ironically) and they are a horrible idea as they are completely implicit and non-obvious in their function.
Using type membership tests + if statements is so 90s. Is Julia also trying to draw in the Java crowd here by trying to be closer to that language?
> <all those other points>
> There are many situations where being forced to know array dimensions at compile time is not helpful or even actively problematic.
It is not that Spiral enforces that dimensions be static. Rather if they are known at compile time then that information is merely propagated forward including through function call boundaries.
You seem to miss what it really means to have first-class types together with staging in a language. Spiral's tensors can be arbitrarily static or dynamic in their dimension and can even allow some dimensions to be static (known at compile time) while the others are dynamic. No friction results from this.
All this does not actually require separate implementations like in Julia. Both Spiral and Julia have Turing complete type systems, but based on this I can conclude that Spiral's is more expressive.
It would be trivial to force it so all the dimensions of a tensor are dynamic, but why would one want to propagate less information during compilation? It is not like dimensions of tensors change that often or arbitrarily like common variables do.
Also let me just state for the record that Spiral is intended to be more than a GPU language. A language with the capabilities of doing it elegantly was my motivation, but Spiral featureset makes it uniquely suited for both very high level and very low level programming.
A language saying it wants to be as fast as C counts for very little, the question is how fast it would be once the code starts being really abstract? This is why the rare one benchmark that currently exists in Spiral's documentation is for parser combinators.
GPU kernels are not a good test bed for language speed because they do not use that many high level features except for the ones needed for tensors. I'd be more concerned with all the scaffolding needed to set up the kernel. And in fact that was one of the majors concerns for me back when I was doing a ML library in F#.
https://www.youtube.com/watch?v=YRoruJRmuLc&t=2537s Jeffrey M. Siskind – The tension between convenience and performance in automatic differentiation
In the field that Julia is aiming for, I'd rather see a benchmark for a CPU based AD library that works on scalars which is a use case in scientific computing. Optimizing this is actually difficult given that none of the mainstream functional languages can do it. I'd expect the same situation as with monads for Julia where the inliner just gives up.
> Of course, we're also cheating since Julia doesn't do type checking and therefore having a highly expressive type system doesn't really cost us much.
I do not know whether the code that was linked in these threads is a representative sample for Julia, but Julia's code to me looks much more like it was written in a static language than Spiral's does.
I think that language wars such as these are necessary in order to come to the truth. It is not that I am being rude on purpose, it is that rudeness in a battle is to be excused and seen as unavoidable. Being right is a process rather than a fact because being able to get it down to a fact is rare and getting the fact to be accepted is even rarer. Cooking a good meal requires flames.
Language semantics have indivisible algorithmic properties that set them apart from each other and hence they cannot ever be equal. The best thing therefore is to enjoy the division. I see this as different than celebrating diversity.
Heh, I've noticed that the setup code tends to come out longer than the actual kernels. It is inside `kernel = cuda`. Spiral is indentation sensitive like Python and F#.
The stuff inside {} is just module creation, think of it like tuples with named fields.
> Do you have any benchmarks for the softmax kernel? If that kernel has optimal performance, it would be quite interesting. If it's sup par, it looks much longer than a simple version.
No, I've yet to actually benchmark it. It really depends on how good of a job NVCC does with the generic sequential reduce kernel. I'll do an in depth analysis when I am done with all the neural network work that I am doing currently which might take a while.
https://github.com/mrakgr/The-Spiral-Language/blob/7ecd30bdf...
You can see how it is implement here for the forward and backward parts.
https://github.com/mrakgr/The-Spiral-Language/blob/7ecd30bdf...
In the actual cost function, it is a bit different since I fuse the forward and the backward parts.
It is a bit ugly because of all the type checking for whether the argument is a dual, but that is all done at compile time.
> https://mikeinnes.github.io/2017/08/24/cudanative.html
This is actually loop unrolling example that I had in mind when I said that Julia needs macros for.
https://github.com/mrakgr/The-Spiral-Language#3-loops-and-ar...
You can in fact get loop unrolling with just standard functions in Spiral. I go into it in the context of this chapter. It is achieved by pattern matching over tuples and recursion.
> If you can make concrete examples of how things can be simpler than this, I'd be delighted to hear them :)
Yes, by replacing meta programming with intensional polymorphism and inlinining guarantees. Also reflection should be done using pattern matching. I think this last one could be taken entirely seriously as it would not involve a replacement of the entire type system.
All the claims in the article about inlining and specialization that make it sound like magic is what in general makes me dubious about languages pretending to be speed kings. Yes, I am aware that GPU kernels do not require optimizers capable of having monads for breakfast and that in the context of GPU programming where they were made they are probably true, but inlining is the sort of thing that matters more the more high level a language is. For very high level languages that desire speed, there isn't much choice but to make them a part of language semantics despite the added burden it puts on the user.
Since we are still at it, I have a question I need to ask.
Recently I've been informed that Julia is capable of GCing GPU memory. If this is fully integrated that would be a major feature which is not possible in say .NET or Racket. I really wanted this in Spiral and could not get it in .NET.
By fully integrated, I mean much like for regular memory for which the GC takes note of the state of the system for when to do collection and defragmentation. If it can only make a thin wrapper with a finalizer (much like in .NET) around an unmanaged resource then it is not a big deal.
Is it fully integrated or is it a wrapper style memory management?
If it is the later, then that is too bad, but I'd suggest to Julia devs that they work on making it fully integrated as it would be a really good feature. Obviously, I can't do it in Spiral as I would need to write my own VM and I have only so many years in my life.
Since I haven't done so and I really should have, if you or anyone else still looking at this thread are curious how GPU programming without macros looks like then take a peek at this:
https://github.com/mrakgr/The-Spiral-Language/blob/3a94ca644...
Comparing Spiral and Julia code that has been linked so far, I feel that no one would ever guess without knowing ahead of time that Spiral is a static language and Julia dynamic based on these examples.
As an example of this in action, here is how layer normalization is implemented using the 'seq_broadcast' kernel which does a sequence of reductions and broadcasts in registers. It is also used to implement softmax for example.
https://github.com/mrakgr/The-Spiral-Language/blob/3a94ca644...
So Julia can definitely do flexible data structure layouts, but looking at your code I am starting to understand why it took an expert such as yourself to show me this example. Had you not told me that the link is to a map kernel I would have difficulty figuring it out on my own.
This shifts my objections from 'Julia cannot do it', to 'macro heavy code is quite hard to grasp'. In Spiral you could have written all of this without the need for macros. Spiral does have (text) macros, but their purpose is to do language interop and not abstraction.
The use of macros is a typical anti-pattern in languages whose type systems are too weak for the problem at hand - I am decently sure that with sufficient macro magic any language that has them would have been capable of performing these same optimizations.
One immediate benefit of a powerful type system that I'd like to point out is that Spiral allows reflection over tuples, which is a lot more elegant solution to a problem of a function having variable arguments than the C-style varargs mechanism that I see Julia using in the examples you've shown me. Also, in Spiral reflection is done using pattern matching rather than if statements.
Another benefit is that Spiral has no need for separate static array machinery. Because literals can be a part of a variable's type and if not, unless explicitly prevented the literals tend to be propagated through function call boundaries (join points). Assuming they are known at compile time the inlining of tensor dimensions in loops is actually the default behavior in Spiral. I've yet to do benchmarking to see how helpful that is, but I am sure it would help reduce register pressure.
Your examples are not quite enough to get me to apologize for my rude behavior, but are enough to get me to shift my views a little so I'll change the offending sentence so it reflects reality more accurately.