We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
finecodex.com
finecodex.com
The page includes the logos of those companies. Is it normal to do that for companies one used to work for?
Sounds fishy
Easier to land investors and customers if you have the "correct" pedigree.
It seems this team is using the old scammy marketing trick of using logos of every company you can claim any possible relationship with as a way of building trust. Plastering giant logos on a product website implies some endorsement or affiliation to most casual readers. It’s not until you read all of the text that you realize this is just a list of companies they worked at.
This is the kind of behavior that earns a sternly worded letter from corporate council. You shouldn’t expect to be able to leave a company and then put their logo on your product page.
Trivial issue because it saves me A LOT of time, but it could be an issue for new people testing it.
I would love to test this approach. Are you guys fine tuning for each codebase?
Their cost is $0.7 per 1M token.
DeepSeek is $0.14 / 1M tokens ( cache miss)
1. Data is used for training
2. Context window is rather small and doesn't fit as well large codebase
I keep saying this over and over in all the content I create, the valu of coding with AI will come from working on big, complex, legacy codebases. Not from flashy demo where you create a to-do app.
For that you need solid models with big context and private inference.
https://api-docs.deepseek.com/quick_start/pricing
Running it locally is quite a bit beyond the scope of being productive while coding with AI.
Beside that 128k is still significantly less than Claude
[1] https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct
Whenever using a model to be more effective as a developer I don't particularly care if the model is open source or closed source.
I would love to use open source models as well, but the convenience to just plug an API against some endpoints in unbeatable.
My email is richard at our website's domain if you'd like to get in touch!
Why would someone advise against it? IMHO that sounds as the end game to me. If it weren't so darn expensive, I'd try this for myself for sure.
If we are approaching diminishing returns it makes more sense to finetune. As the recent advances seem to happen by throwing more compute to CoT etc maybe the time is close or has already come.
In general I'd suggest trying this first:
- Large context: use large context models to load relevant files. It can pickup your style/tool choices fine this way without fine tuning. I'm usually manually inserting files into context, but a great RAG solution would be ideal.
- Project specific instructions (like .cursorrules): tell it specific things you want. I tell it preferred test tools/strategies/styles.
I am curious to see more detailed evals here, but the claims are too high level to really dive into.
In generally: I love fine tuning for more specific/repeatable tasks. I even have my own fine-tuning platform (https://github.com/Kiln-AI/Kiln). However coding is very broad. Good use case for foundation large models with smart use of context.
Kudos to the founders for shipping. I do think this kind of functionality will become very rapidly commoditized though. But then, I suppose people said the same thing about Dropbox.
0. Is this a scam? No. We're very early (started <1 month ago) so our landing page is to validate our concept, gather initial feedback and start conversation on what we can build that would be most applicable. We'll add more details and benchmarks to the website.
1. Company logos. You're right. We're using our work experience as a credibility signal because at this stage that is our main selling point. We'll replace logos with concrete results as we develop.
2. Team. We're 2 software engineers and 1 AI researcher: - I was an AI product engineer at Asana. https://linkedin.com/in/samatd - Denis was a tech lead at a unicorn startup. https://x.com/karpenoid - Our third co-founder works at Anthropic and was previously at OpenAI. Since he is still at Anthropic and planning to leave soon for the startup, I can share his details privately.
3. Claims and transparency. Our "4.2x Sonnet-3.5 accuracy" is an initial estimate from a locally fine-tuned model. Actual results may vary - a small app might not see big improvements, but we believe larger, private enterprise projects could see significant gains. We plan to publish our fine-tuned model so others can verify the results.
4. Competition from LLM providers. Fine-tuning requires complex data cleanup and setup. Enterprise projects have fragmented data, making automation challenging for big providers like OpenAI.
Appreciate the feedback! If you want to chat more 1-1, happy to discuss at hi@finecodex.com Samat.
2024: Our tiny model blah blah blah beats Claude!
2025: Our tiny model blah blah blah beats ???
Edit: OK, right it's olama, so I assume you can download your own model. (Assuming it's downloadable?)
I think openAI already offers fine-tuning with custom data for some of their models, but maybe not specific to coding tasks.
Is this for completions, patches, or new files?