HNHacker News
TopNewBestAskShowJobs

RuiWang0811

2 karma · joined March 2, 2025

Ex-Quant turned founder. Interested in all AI + trading related things.
submissionscomments
RuiWang0811··on Research Harness from First Principles
TLDR: we built Oh My Quant, a harness designed for research from first principles -> fast & cheap experiment running, monitoring and storing. Our harness with Opus is top 1 on MLE leaderboard.

The point is to let an agent spawn, run, and track a lot of hypotheses and experiments in parallel without confusing itself.

Most harnesses right now are for coding, which is fine for SWE but not ideal for research. Good research means shipping many hypotheses, experiments, and validations, instead of static PRs. We used Codex and Claude Code for a while, but they have poor research 'taste' and are bad at running and comparing multiple experiments.

We ended up designing Oh My Quant from scratch for our own research. The harness is built around hypothesis → experiment → evidence, with a few simple principles: no hardcoded workflows, main agent decides which subagents to spawn; cheap and fast experiment design and tracking; harness handles all tedious things like parsing experiment metadata, run and monitor experiment, store result metrics - and leaves the agent's entire context window free to do the intelligent judgement calls for research design.

RuiWang0811··on Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
there are loads of env businesses for coding tasks, enterprise tasks, computer use etc etc. we offer different envs from a niche industry, which just so happens to be a very hard data science task & where the data doesn't saturate. In order to build these you'd need niche expert knowledge.
RuiWang0811··on Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
We have experiments showing that agents at least can learn from the environment by overfitting on train. But we do not yet have full post train runs, mainly due to time. But follow our blog/X where we’ll regularly update our research
RuiWang0811··on Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
We use real historical market data for the environments. There is no parametric modelling involved.

The decay property refers to alpha that we give the agent for trade in the env - they are generated as tools.

RuiWang0811··on Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
No we sell our own research to AI labs as RL envs. Realistic RL envs grows in demand as labs seek better data train better models.

It’s a complimentary business. Simply put: we sell envs to labs, labs make better models, firms buy these models to make more profit. Everybody wins.

Yeah a competitor for us would be fellow quants doing the same thing. But even then every quant firm trades differently (and good ones all make money) so envs can still be sufficiently different.

RuiWang0811··on Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
languagelearner, I think you need to spend more time learning languages
RuiWang0811··on Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
this seems to be a common misconception, our envs use market data, but the goal is not (only) trading. Market data just happens to be a good source of hard data science tasks.

Re trading: I’d argue there is no such thing as solving investment nor is there “the one profitable strategy”. Every decision from personal risk appetite to trading horizon changes what is the optimal strategy for you and there are multiple strategies that make money.

Also note that even the most profitable alphas are no crystal balls. Someone else mentioned 5% correlation to future return - depending on horizon and data such level of correlation can make 9 figure PnL and is by no means easy to achieve

RuiWang0811··on Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
we do affine transformations of the data, so all return/ pnl measures are still the same as with untransformed data. The transformation doesn’t change the conditional distribution of the data, which is what alphas ultimately measure
RuiWang0811··on Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
not sure about your background, the trace shows the feature engineering the LLMs did
RuiWang0811··on Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
cofounder here - LLMs can do some model training, they train on ML competition data after all. But they do struggle with low signal to noise ratio of market data. But that’s exactly what our environments will teach.
RuiWang0811··on Launch HN: Tokenless (YC S26) – Automatic model switching to save money
Does this mean you have to retrain routing rules every time a new model gets released? I imagine since the price/token (or rather the amount of work that can be done per token) does not monotonically increase with new models, that the routing logic has to change all the time
RuiWang0811··on [dead]
With all the benchmaxxing happening, current evals are oversaturating and become meaningless for model comparison. I think there is only one test that can't be gamed: let them trade in real markets.

Markets are self-improving, as models trade better, the market inefficiencies disappear, so trading well gets harder the better you get. It's a never-ending hill to climb.

I ran SOTA reasoning LLMs on it. TL;DR: they suck, and reasoning doesn't help. No model came close to a simple static benchmark (experienced human Quant chose parameters);More reasoning ≠ better trading; Fun fact: when losing, LLMs trade less rather than smarter.

Why quant trading is a great intelligence test (minimal jargon): a) requires ML research that generalizes OOS (alpha); b) needs robustness to regime shifts and new rules; c) trains long-horizon planning under tradeoffs (how to trade now if AAPL is +2% tomorrow but −5% the day after). And d) it's self-correcting: alphas decay, trading too much gets you adversely selected, and trading well makes the market more efficient — so it never stops being hard.

Background: Ex-quants from G-Research (ML algo trading) & TransMarketGroup (Crypto options); + LLM inference kernels at Etched.

Please poke holes in the setup - where does "markets as an eval" break down? Especially keen on views from outside quant.