HNHacker News
TopNewBestAskShowJobs

olokobayusuf

174 karma · joined March 18, 2018

submissionscomments
olokobayusuf··on Making PyTorch –> Qualcomm NPUs less treacherous
There are over 2.5 billion Qualcomm processors in the world today (PC, mobile, automotive, etc). But the process for bringing AI models to run on Qcom processors is a (massive) pain. Their 2GB+ SDK is an encyclopedia's worth of information needed to deploy correctly.

We're working to make Qualcomm NPUs a first-class citizen for deployment from PyTorch. Devs can write a Python function that runs a PyTorch model, then use our `@compile` decorator to transpile the model to a Qcom-specific C++ implementation (DLC) which compiles to a self-contained shared library.

The Qualcomm NPUs are fast. 1.8x faster than ONNXRuntime. See the link above.

olokobayusuf··on Kubetorch – For RL and ML on Kubernetes
Congrats on the launch!
olokobayusuf··on Launch HN: LlamaFarm (YC W22) – Open-source framework for distributed AI
We should collab! We prefer to be the underlying infrastructure behind the scenes, and have a pretty holistic approach towards hardware coverage and performance optimization.

Read more:

- https://blog.codingconfessions.com/p/compiling-python-to-run... - https://docs.muna.ai/predictors/ai#inference-backends

olokobayusuf··on Launch HN: LlamaFarm (YC W22) – Open-source framework for distributed AI
We're building something closer to this at Muna: https://docs.muna.ai . Check us out and let me know what you think!
olokobayusuf··on Launch HN: LlamaFarm (YC W22) – Open-source framework for distributed AI
This is super interesting! I'm the founder of Muna (https://docs.muna.ai) with much of the same underlying philosophy, but a different approach:

We're building a general purpose compiler for Python. Once compiled, developers can deploy across Android, iOS, Linux, macOS, Web (wasm), and Windows in as little as two lines of code.

Congrats on the launch!

olokobayusuf··on Python developers are embracing type hints
Yup that's true. We do benefit from massive efficiencies though, thanks to LLM codegen.
olokobayusuf··on Python developers are embracing type hints
Our primary use case is cross-platform AI inference (unsurprising), and for that use case we're already in production by startups to larger co's.

It's kind of funny: our compiler currently doesn't support classes, but we support many kinds of AI models (vision, text generation, TTS). This is mainly because math, tensor, and AI libraries are almost always written with a functional paradigm.

Business plan is simple: we charge per endpoint that downloads and executes the compiled binary. In the AI world, this removes a large multiplier in cost structure (paying per token). Beyond that, we help co's find, eval, deploy, and optimize models (more enterprise-y).

olokobayusuf··on Python developers are embracing type hints
I'm founding a company that is building an AOT compiler for Python (Python -> C++ -> object code) and it works by propagating type information through a Python function. That type propagation process is seeded by type hints on the function that gets compiled:

https://blog.codingconfessions.com/i/174257095/lowering-to-c...

olokobayusuf··on Ask HN: Who is hiring? (June 2025)
Function (https://fxn.ai) | Remote (US)

We're building native code generation for AI developers. We generate high-performance C++/Rust to power open-source and on-device AI for our customers. We have customers ranging from early stage startups to the Fortune 1000.

You'll be:

1. Writing open-source Python functions that run popular vision models and LLMs; or

2. Writing high-performance C++ and Rust code that targets different accelerators (CUDA, Metal, etc); or

3. Writing parts of our Python-to-C++ compiler in support of (1) and (2); or

4. Some combination thereof.

Join the party: Email us at stdin@fxn.ai or apply at https://app.dover.com/jobs/fxn.

No recruiters; no visa sponsorship (yet). We prize demonstrated curiosity and impact over everything else.

olokobayusuf··on Show HN: Powering React with Python (WASM)
Link to original article that kickstarted all of this: https://medium.com/@eugeniyoz/powering-angular-with-rust-was...
olokobayusuf··on Show HN: Python at the Speed of Rust
When we trace Python code, devs have to explicitly opt-in dependency modules to tracing. Specifically, the `@compile` decorator has a `trace_modules` parameter which is a `list[types.ModuleType]`.

With this in place, when we trace through a dev's function, a given function call is considered a leaf node unless the function's containing module is in `trace_modules`. This covers the Python stdlib.

We then take all leaf nodes, lookup their equivalent implementation in native code (written and validated by us), and use that native implementation for codegen.

We don't interact with the GIL. And we keep track of what is unsupported so far: https://docs.fxn.ai/predictors/requirements#language-coverag...

olokobayusuf··on Show HN: Python at the Speed of Rust
Not entirely sure what you mean by having to deal with C-API and having a language spec.

We're also not competing with LLMs at all--we use LLMs for said conversion (under strict verification requirements).

olokobayusuf··on Show HN: Python at the Speed of Rust
Actually we're currently implementing Numpy (and PyTorch) support, and will cover a few other core scientific computing libraries like scipy. See docs: https://docs.fxn.ai/predictors/requirements#library-coverage
olokobayusuf··on Show HN: Python at the Speed of Rust
Not quite.

First, Function is designed to be truly cross-platform but libraries like Numpy aren't compiled for say WebAssembly.

Second, the native libraries are usually built around CPython interop (i.e. the C API expects to interact with the CPython interpreter). Function does not (and will never) have a CPython interpreter (we generate full AOT compiled code).

olokobayusuf··on Show HN: Python at the Speed of Rust
Spot on!

The majority of the innovation here is in building enough rails (specifically around lowering Python's language features to native code) so that LLM codegen can help you transform any Python code into equivalent native code (C++ and Rust in our case).

olokobayusuf··on Show HN: Python at the Speed of Rust
I think a more pedantic way to describe what I mean is:

"What if we could compile Python into raw native code *without having a Python interpreter*?"

The key distinguishing feature of this compiler is being able to make standalone, cross-platform native binaries from Python code. Numba will fallback to using the Python interpreter for code that it can't jit.

olokobayusuf··on Show HN: Python at the Speed of Rust
Yes, we upload user code to a cloud sandbox in order to run our symbolic tracing and code generation algorithm.

Beyond that, we also compile the generated native code in the cloud so that devs don't have to have a cross-compiler installed on their system (i.e. Clang, Xcode, MSVC, and so on).

olokobayusuf··on Show HN: Python at the Speed of Rust
Way ahead of you: https://github.com/olokobayusuf/python-vs-rust/blob/main/Car...

I've clarified that this is not designed to be a rigorous benchmark. We've got rigorous benchmarks coming for image processing and CNN inference. I'll reply with the image processing example benchmark this week.

olokobayusuf··on Show HN: Python at the Speed of Rust
More serious reply: us and Mojo have similar visions of where the world is going. The key difference between us is that Mojo is itself a new programming language. Sure, it supports Python, but it doesn't actually compile Python. It simply delegates all interactions with Python code to the CPython interpreter: https://docs.modular.com/mojo/why-mojo/#compatibility-with-p...

But beyond that, our goal is meeting devs where they are. This means that beyond just compiling Python code, we provide SDKs for different frameworks (JavaScript, Kotlin, Swift, React Native, Unity, etc) that devs can use to run these functions within their applications, in as little as two lines of code.

We're very (very) focused on developers shipping products that use Function. We're already embedded in web apps, apps on the App Store, Play Store, and other places.

olokobayusuf··on Show HN: Python at the Speed of Rust
We use LLVM!

We're not translating Python directly to LLVM IR (I think I've seen other projects do this). We translate Python to C++/Rust first, where we have rigorous unit tests for every operation we support translating. We then use LLVM for downstream compilation to object code.

Here's some more context: https://docs.fxn.ai/predictors/compiler

olokobayusuf··on Show HN: Python at the Speed of Rust
There's actually no overhead of "still using Python" because we don't use Python. The overhead in Function exists solely because we have a bunch of sugar added to create a unified interface for calling different kinds of functions (i.e. generator functions, etc).

You don't have to take my word for it: you can pull the C++ source code that we generate from a given Python function and inspect it yourself: https://docs.fxn.ai/predictors/compiler#compiling-binaries

olokobayusuf··on Show HN: Python at the Speed of Rust
https://x.com/OlokobaYusuf/status/1908538983874810303
olokobayusuf··on Show HN: Python at the Speed of Rust
We're focused specifically on on-device AI inference (and related compute-bound algorithms, like computer vision or scientific computing).

We want devs to find or develop inference code (always in Python); decorate it with Function's `@compile`; compile it; and run natively in their Android, iOS, macOS, Linux, WebAssembly, or Windows applications.

We won't be supporting other use cases like web servers (no Django or FastAPI).

olokobayusuf··on Show HN: Python at the Speed of Rust
Yup but you're skilled enough to write--and more importantly, maintain--the required C/C++ code. Most devs and companies we talk to just want to make something people want; they don't care for the added complexity of writing and maintaining native code.

The way I like to think about this is how much more code got written when "high level" languages like C came onto the scene, at a time when Assembly was the default. Writing Python is way (way way) easier and faster than any of the lower-level languages--no pointers, no borrow checker!

olokobayusuf··on Show HN: Python at the Speed of Rust
No need to be snarky. The choice to default to FP32 is inspired by the fact that most typical use cases don't need double-precision (GPU shader languages and game engines do this all the time). This in turn allows us to vectorize code for 2x throughput compared to using FP64. We're gonna add a flag to change the default floating-point precision for devs who need extra precision.

See docs: https://docs.fxn.ai/predictors/requirements#floating-point-v...

olokobayusuf··on Show HN: Python at the Speed of Rust
Thank you <3
olokobayusuf··on Show HN: Python at the Speed of Rust
Spot on! We're starting from specific functions for now, and will expand the scope if/when feasible. The challenges exist across two axes: language coverage and library coverage.

For language coverage, Function currently doesn't support classes (on the roadmap) or lambda expressions (much harder). But these are the main limitations.

For library coverage, this is where things get incredibly hairy. If a library isn't pure Python (i.e. most of the important libs), we can't compile them. We instead have to reimplement the library's functionality. For now, we plan on using LLMs to automate this as much as possible.

olokobayusuf··on Show HN: Python at the Speed of Rust
It's marginal. The difference exists because Function adds syntactic sugar to make the developer experience of calling the compiled functions easy and smooth. The Function call isn't a direct call; whereas the Rust one is.

It's possible to hack Function to perform a direct call and avoid (most of) the overhead, but in 99.99% of use cases this won't matter cos most processing time will be spent in the body of the function, not the scaffolding that Function uses.

olokobayusuf··on Show HN: Python at the Speed of Rust
Also worth noting that we're better positioned for the coming wave of LLMs writing the vast majority of code in production. Function is designed specifically for LLMs to do the transpilation step (Function simply provides rails to stack these unitary operations into user-defined functions).

Mojo has a cold-start problem here, cos there isn't enough Mojo code in the wild for LLMs to be great at writing it in volume (compare this to C++ and Rust).

olokobayusuf··on Show HN: Python at the Speed of Rust
I think us and Modular have incredibly similar visions of the end state of the world. The main difference is that they require devs to learn a new programming language (Mojo), whereas Function is designed to meet devs exactly where they are--not an inch away.

More fundamental than that is that Mojo has somewhat of a built-in assumption (I'm a bit more skeptical) that they can outperform silicon OEMs like Nvidia using MLIR (Mojo is designed as a front-end to MLIR). We have a much more conservative view: we'll rely on silicon OEMs to give us libraries we can use to accelerate things like inference; and we'll provide devs the ability to inspect and modify the generated code (C++ atm, Rust soon).

TL;DR: Devs don't have to learn anything with Function. Devs have to learn Mojo and MLIR to use Modular's offerings properly.

Page 1 of 2Next →