HNHacker News
TopNewBestAskShowJobs

samatdav

4 karma · joined October 29, 2020

hi@finecodex.com
submissionscomments
samatdav··on Ask HN: Please Review My Startup Chatty AI – Operator for Your Website
cool! So the user does not need to use any additional tools?
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Haha, yes it is a pattern. However, the claim here is that "our tiny model beats best model" is applicable for highly specific tasks.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Yes, you can download and host the fine-tuned open-source model like Llama. The fine-tuning is easy once you have the data, but gathering and cleaning data is challenging. There are also optimizations like upsampling and distillation that could improve the quality of the resulting model. We had 40 engineers at the Asana AI org and never did the fine-tuning because it is not easy.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Thank you, we will!:) This was a quick landing page for us to start the conversation and gather feedback. We are trying to make sure we are not building something that nobody needs.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
We used a single file for the context. It is a cherry-picked example, you are right. I wanted to demonstrate a simple visual change that our model did correctly unlike Sonnet-3.5. Since we are just getting started, we don't have many features like making changes across multiple files in the code editor so it would be harder to demo. Our premise is that a smaller fine-tuned works better than a large, general-purpose SOTA model. We plan to share more metrics and data in the future.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Good point, I agree, we haven't shared enough details. Since we are very early, we only got high level results and want to get feedback on what direction would be most applicable and useful. We plan to add more metrics and data to the website in the future and also want to publicly host a fine-tuned model for anyone to try and see.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
I agree. Our local early results were promising were a higher percentage of code change requests produced a functionally correct output. We will post more metrics and data in the future.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Not yet, but we plan to publicly host a fine-tuned model so anyone can try.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
We run a set of change requests on the discourse repo. Good point, we plan to publish more detailed testing benchmarks and metrics on the website.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Good point, we plan to publish more benchmarks and also publicly host a model for anyone to try. We think Llama is a good option but as we progress we will test other open source models too like deepseek.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Thank you! Will email you within a couple of days:)
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Yes, we fine-tune for each codebase. Now we are focusing on larger enterprise codebases that would: 1. benefit from the fine-tuning the most. 2. have the budget to pay us for the service. For smaller projects that are price-sensitive we are probably not a good fit at this point.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Thank you for the idea! We are also considering upsampling and distillation. But on high level, correctly setting up the data for simple fine-tuning can already produce great results.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
I agree, we plan to publish more benchmarks and metrics. We also want to publicly host our fine-tuned model for one of the open-source repos so that people can try themselves agains SOTA models.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Good point, we should provide more detailed metrics. Since we are very early, we focus on the main metric in our view: higher accuracy of changes to be more practically usable. We will do more testing on overfitting and how the model performance on different types of tasks. On high level we believe in the idea of "a well fine-tuned model should be much better than a large general model". But we need more metrics, I agree.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Thank you for the suggestion, we will take a look!
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Looks like a great repo to try the fine-tuning! I will email you, thanks!
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Could be done in the future. Our current focus is highest accuracy. But there are no limitations on the models - just would depend on user preference of size/performance tradeoff.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
I agree, we need to post more data. Since we are very early (<1 month) we just shared the initial results. Discourse repo was just a good option since it is a big public repo that could benefit from fine-tuning. We plan to add more benchmarks to the website as we progress.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
I understand the concern but we don't need anyone's IP. Unfortunately, it is hard to provide fine-tuning solution without access to the codebase. We just think that using a large general-purpose model for a highly specific codebase with a lot of internal frameworks is not the best solution and want to try to improve it.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
At Asana we did not do any fine-tuning because it was too complicated even for our AI org of 40 engineers. We believe we can do it by setting up and cleaning data correctly.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Good point! We are just very early and our experience is our main selling point. We plan to remove it.
samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Hi HN! I'm Samat, the co-founder from the video. Thank you for the critical feedback, great points.

0. Is this a scam? No. We're very early (started <1 month ago) so our landing page is to validate our concept, gather initial feedback and start conversation on what we can build that would be most applicable. We'll add more details and benchmarks to the website.

1. Company logos. You're right. We're using our work experience as a credibility signal because at this stage that is our main selling point. We'll replace logos with concrete results as we develop.

2. Team. We're 2 software engineers and 1 AI researcher: - I was an AI product engineer at Asana. https://linkedin.com/in/samatd - Denis was a tech lead at a unicorn startup. https://x.com/karpenoid - Our third co-founder works at Anthropic and was previously at OpenAI. Since he is still at Anthropic and planning to leave soon for the startup, I can share his details privately.

3. Claims and transparency. Our "4.2x Sonnet-3.5 accuracy" is an initial estimate from a locally fine-tuned model. Actual results may vary - a small app might not see big improvements, but we believe larger, private enterprise projects could see significant gains. We plan to publish our fine-tuned model so others can verify the results.

4. Competition from LLM providers. Fine-tuning requires complex data cleanup and setup. Enterprise projects have fragmented data, making automation challenging for big providers like OpenAI.

Appreciate the feedback! If you want to chat more 1-1, happy to discuss at hi@finecodex.com Samat.

samatdav··on We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
Hi! Currently we generate a whole diff (like cmd+shift+k in Cursor). But plan to add there rest soon!:)
samatdav··on Show HN: Mode – Your Personal AI Code Copilot for VS Code
Exciting - will give it a try!:)