HNHacker News
TopNewBestAskShowJobs

mikepapadim

80 karma · joined March 9, 2020

submissionscomments
mikepapadim··on Claude Cowork found me a flat to rent in London in just 5 days
Twice a day, Claude Cowork searched SpareRoom, OpenRent, Rightmove and Zoopla, filtered out student flats and 3+ bed houses, wrote personalised outreach messages forevery good listing, and emailed me a digest I could act on from my phone.

As a result, I managed to put down a deposit for a decent 1 bed flat in London in just under a week for the price people asking for a room!

Totally skipped all the manual searching of new ads etc.

I create a repo so others can take ideas or even reuse it.

Just adapt your preferences

Just with Claude cowork + Claude in Chrome + Gmail MCP

mikepapadim··on 5 Open-Source Projects Running Transformers on CPUs to GPUs in Pure-Modern Java
- [Llama3.java] https://github.com/mukel/llama3.java - [Gemma4.java] https://github.com/mukel/gemma4.java - [Jlama] https://github.com/tjake/Jlama - [GPULlama3.java https://github.com/beehive-lab/GPULlama3.java - [Qxotic] https://github.com/qxoticai/qxotic - [TornadoVM]https://github.com/beehive-lab/TornadoVM
mikepapadim··on 5 Open-Source Projects Running Transformers on CPUs to GPUs in Pure-Modern Java
- [Llama3.java](https://github.com/mukel/llama3.java) - [Gemma4.java](https://github.com/mukel/gemma4.java) - [Jlama](https://github.com/tjake/Jlama) - [GPULlama3.java](https://github.com/beehive-lab/GPULlama3.java) - [Qxotic](https://github.com/qxoticai/qxotic) - [TornadoVM](https://github.com/beehive-lab/TornadoVM)
mikepapadim··on [dead]
I'm exploring an idea to reduce GPU dispatch overhead in a Java-based runtime, TornadoVM, which executes compute operations from TornadoVM bytecodes.
mikepapadim··on Show HN: GPULlama3.java Llama Compilied to PTX/OpenCL Now Integrated in Quarkus
https://github.com/beehive-lab/TornadoVM/pull/732 https://github.com/beehive-lab/TornadoVM/pull/313
mikepapadim··on Show HN: GPULlama3.java Llama Compilied to PTX/OpenCL Now Integrated in Quarkus
Yes, when you use the PTX backend it supports Tensor Cores.It has also implementation for flash attention. You can also write your own kernels, have a look here: https://github.com/beehive-lab/GPULlama3.java/blob/main/src/... https://github.com/beehive-lab/GPULlama3.java/blob/main/src/...
mikepapadim··on Show HN: GPULlama3.java Llama Compilied to PTX/OpenCL Now Integrated in Quarkus
https://github.com/beehive-lab/GPULlama3.java
mikepapadim··on [dead]
Highlights Simpler execution via Java argfiles Improved performance on FP16/Int8 LLM Inference on hashtag#Nvidia GPUs Extended reduced precision type support for GPUs (Int8, fp16) Zero-copy object support through project Panama Support for compressed oops on modern JVMs New cross-platform SDK distribution (soon hashtag#SDKMAN! https://lnkd.in/d8pGHYy5) Official TornadoVM dependencies now published on Maven Central. (https://lnkd.in/dDRZj8ru)
mikepapadim··on [dead]
https://github.com/beehive-lab/TornadoVM
mikepapadim··on Ask HN: Who wants to be hired? (July 2025)

  Location: UK
  Remote: Yes
  Willing to relocate: No
  Technologies: Java, C++, Python, Cuda, OpenCL, Docker, ONNXRT,Git, Apache TVM, Compilers, GPUs
  Résumé/CV:https://github.com/mikepapadim && https://www.linkedin.com/in/michalis-papadimitriou/
  Email:mpapadimitriou92 [ΑΤ] gmail.com
mikepapadim··on [dead]
https://github.com/beehive-lab/GPULlama3.java

We took Llama3.java and we ported TornadoVM to enable GPU code generation. Apparrently, the first beta version runs on Nnvidia GPUs, while getting a bit more than 100 toks/sec for 3B model on FP16.

All the inference code offloaded to the GPU is in pure-Java just by using the TornadoVM apis to express the computation.

Runs Llama3 and Mistral models in GGUF format.

It is fully open-sourced, so give it a try. It currently run on Nvidia GPUs (OpenCL & PTX), Apple Silicon GPUs (OpenCL), and Intel GPUs and Integrated Graphics (OpenCL).

mikepapadim··on GPULlama3.java – Llama3.java on Steroids
Java to OpenCL and PTX inference of Llama3 through TornadoVM
mikepapadim··on Llama Deck:CLI for running multiple language implementations of LLM inference
Llama Deck is a command-line tool for quickly managing and experimenting with multiple versions of llama inference implementations. It can help you quickly filter and download different llama implementations and llama2-like transformer-based LLM models. We also provide some Docker images based on some implementations, which can be easily deploy and run through our tool.
mikepapadim··on Multifaceted Memory Analysis of Java Benchmarks with NUMAProfiler and PerfUtil [pdf]
A comprehensive analysis of the memory behavior of 30 Dacapo and Renaissance Java applications using a dual profiling methodology with NUMAProfiler and PerfUtil in MaxineVM, identifying various memory pressures and JVM impacts.
mikepapadim··on Radicle: Open-Source, Peer-to-Peer, GitHub Alternative
How is this related to the $RAD coin?
mikepapadim··on Ask HN: Do you know any new llama2.c implementations not mentioned in the repo
https://github.com/mikepapadim/llama-shepherd-cli
mikepapadim··on Llama2-shepherd a CLI tool to install multiple implementations of the llama2
most of the models support the tinyllamas, regarding gguf/ggml and safetensors each implementation has its own model importers, so there is not guarantee that all types can be consumed by all implementations
mikepapadim··on Llama2-shepherd a CLI tool to install multiple implementations of the llama2
Not yet, this is the end-goal of this repo, to be able to do this kind of perf evaluation.
mikepapadim··on Llama2-shepherd a CLI tool to install multiple implementations of the llama2
hello, I am planning to add also a runner option and a benchmarking option among different implementations. This is just an MVP version while trying to keep track of all llama2 implementations as the ones in the original repo is a bit outdated
mikepapadim··on TornadoVM: Running Java on GPUs and FPGAs
Indeed aparapi worked by converting java to OpenCL. However, exposed to the user many aspects of GPU programming such as thread indexing (e.g global ids) and memory allocation for specific optimizations (e.g local memory). In the case of TornadoVM these aspects handled by the compiler and the runtime.