I've been playing around with some automated code review tools recently, and it's surprising how often they flag things that are technically correct but just... unusual. Style matters, especially for maintainability.
3,064 karma · joined February 23, 2008
I've been playing around with some automated code review tools recently, and it's surprising how often they flag things that are technically correct but just... unusual. Style matters, especially for maintainability.
Start with understanding parallel computing concepts and how GPUs are structured for it. Optimization is key - learn about memory access patterns, thread management, and how to profile your code to find bottlenecks. There are tons of great resources online, and NVIDIA's own documentation is surprisingly good.
As for the data engineering side, tbh, it's tougher to get into MLE without ML knowledge. However, focusing on the data pipeline, feature engineering, and data quality aspects for ML projects might be
Traditional classifier adding a new class:
- Requires full retraining (~30-60 minutes on typical dataset)
- Needs all historical data
- Uses 2-3x more memory during training
This approach:
- Adds new class in seconds
- Needs only examples of new class
- Memory usage stays constant
- Maintains 95%+ accuracy on existing classes
The code is well-documented and tested. I've included detailed examples showing:
- Batch processing for large datasets
- Multi-language support
- Model persistence
- Custom transformer models
Happy to share more details about the architecture or specific implementation challenges!
The core architecture combines a transformer model for embeddings with a prototype memory system and an adaptive neural head.
When adding new classes, it uses Elastic Weight Consolidation (EWC) to preserve performance on existing classes while learning new ones. This prevents the common problem of catastrophic forgetting.
The prototype memory system maintains class prototypes that get updated efficiently as new examples are added, making it memory-efficient even with large datasets.
All state (prototypes, examples, neural weights) can be saved and loaded, making it easy to deploy and update models in production.
The library is built on PyTorch and integrates with the HuggingFace ecosystem. It's tested with Python 3.8+ and requires minimal dependencies.
Let me know if you'd like me to explain any part in more detail!
The models have gotten much better at generating them with just the prompt. I have not implemented strict support for structured output or JSON generation yet. The response from the proxy are all raw text responses.
One way would be to just apply outlines or some library as a plugin to enable structured outputs.
optillm is an OpenAI API compatible optimizing inference proxy which implements several state-of-the-art techniques that can improve the accuracy and performance of LLMs. The current focus is on implementing techniques that improve reasoning over coding, logical and mathematical queries. It is possible to beat the frontier models using these techniques across diverse tasks by doing additional compute at inference time.
Here is an example of how it is very useful especially for newer libraries.
Recently, we have added support for plugins that enable capabilities like memory, privacy and code execution to optillm. Plugins are just python scripts that you can also write yourself, optillm would then load them at start from the directory.
You can now also combine the plugins and techniques using & and | operators. E.g. We recently evaluated the new FRAMES benchmark from Google. Using a combination of plugins and techniques (we used readurls&memory-gpt-4o-mini) we were able to get 65.7% accuracy on the benchmark which is very close to what Google reported in their paper with Gemini Flash 1.5 (66.5) which has a context length that is almost 10 times that of gpt-4o-mini.
In my experience it does work quite well, but we probably need different techniques for different tasks.