HNHacker News
TopNewBestAskShowJobs

3d27

20 karma · joined November 9, 2023

submissionscomments

How to evaluate multi-turn LLM chatbots

confident-ai.com·3 pts·3d27·
0

We wrote a comprehensive guide on LLM security

confident-ai.com·1 pts·3d27·
0

How to generate synthetic data using SOTA data evolution methods

confident-ai.com·1 pts·3d27·
0

How to build your own LLM evaluation framework

confident-ai.com·2 pts·3d27·
0

Overview of All Major LLM Benchmarks

confident-ai.com·1 pts·3d27·
0

I wrote an article about everything I know about LLM metrics

confident-ai.com·2 pts·3d27·
1

Best practices I learnt from helping health tech enterprise test LLMs

confident-ai.com·1 pts·3d27·
0

Am I too needy? From a data science perspective

medium.com·1 pts·3d27·
1

Best Practices for Unit Testing RAG Systems in Prod

confident-ai.com·4 pts·3d27·
0

Tried Apple's Vision Pros, would not recommend it

theverge.com·2 pts·3d27·
0

Everything I know about LLM evaluation metrics

confident-ai.com·7 pts·3d27·
0

Google 2024 Layoffs on a rolling-basis

1 pts·3d27·
0

Meta Going All in on GenAI

datacenterdynamics.com·3 pts·3d27·
2

I used QAG to implement an LLM text summarization evals

confident-ai.com·3 pts·3d27·
0

I found a way to code like Shakespear

shakespearelang.com·1 pts·3d27·
1

I implemented 12+ LLM evaluation metrics so you don't have to

old.reddit.com·4 pts·3d27·
1

AI Makes Commercial Masterpiece [video]

youtube.com·2 pts·3d27·
0

Show HN: I implemented evals metrics for LLMs that runs locally on your machine

github.com·22 pts·3d27·
3

Overcoming the biggest barrier to practical quantum computers

breakingdefense.com·1 pts·3d27·
0

Google's new model is good but the demo's not reproducible in Bard

boingboing.net·1 pts·3d27·
0

What Is RAG? (With Examples)

confident-ai.com·1 pts·3d27·
0

Found this weird programming language

en.wikipedia.org·2 pts·3d27·
0