HNHacker News
TopNewBestAskShowJobs

stephantul

711 karma · joined September 3, 2024

NLP engineer
submissionscomments

Hiding a Prompt in a Tokenizer

stephantul.github.io·2 pts·stephantul·
0

From Chesterton's fence to Chesterton's gap

stephantul.github.io·86 pts·stephantul·
56

Why scikit learn's fit transform is probably not for you

stephantul.github.io·1 pts·stephantul·
0

Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

github.com·8 pts·stephantul·
0

Show HN: Semble – Fast code search for agents with near-transformer accuracy

github.com·7 pts·stephantul·
0

Show HN: Skeletoken, a Python package for editing model tokenizers

github.com·1 pts·stephantul·
0

Show HN: PyNIFE. 400-900× speedup for embedding-based retrieval pipelines

github.com·2 pts·stephantul·
0

Show HN: Skeletoken, a Package for Editing Tokenizers

github.com·1 pts·stephantul·
0

Turning any tokenizer into a greedy one

stephantul.github.io·2 pts·stephantul·
1

Decasing Transformers for Fun

stephantul.github.io·3 pts·stephantul·
1

Model2Vec as a Fasttext Alternative

minish.ai·5 pts·stephantul·
1

Using overloads to handle union return types in Python

stephantul.github.io·1 pts·stephantul·
1

Ask HN: Favourite resources for learning programming type theory?

6 pts·stephantul·
8

Evaluating ML classifiers using relative error instead of absolute accuracy

stephantul.github.io·1 pts·stephantul·
0

Defeat stringly typing without making your users unhappy

stephantul.github.io·2 pts·stephantul·
0

Distilling ModernBERT into a static model doesn't work

minishlab.github.io·5 pts·stephantul·
3

Show HN: SemHash – Fast Semantic Text Deduplication for Cleaner Datasets

github.com·6 pts·stephantul·
0

Train faster static embedding models with sentence transformers

huggingface.co·52 pts·stephantul·
1

Semhash: Fast deduplication and dataset multitool in Python

minishlab.github.io·3 pts·stephantul·
1

Model2Vec: Make sentence transformers 500x faster on CPU, 15x smaller

huggingface.co·5 pts·stephantul·
0

Show HN: Model2Vec: make sentence transformers 500x faster on CPU, 15x smaller

github.com·9 pts·stephantul·
2

Show HN: Model2Vec: make sentence transformers 500x faster on CPU, 15x smaller

github.com·6 pts·stephantul·
2