DatBench fixes VLM evals: 70% blindly solvable, 42% mislabeled, 35% prod gapdatologyai.com·5 pts·hurrycane·0
Luxical: Lexical-Dense Embeddings for Web-Scale Data Curation (3×–100× Faster)datologyai.com·3 pts·hurrycane·0
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-Scale Pretrainingblog.datologyai.com·1 pts·hurrycane·0
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-Scale Pretrainingblog.datologyai.com·3 pts·hurrycane·0
Image-Text Curation for 1B+ Data: Faster, Better, Smaller Clip Modelsdatologyai.com·12 pts·hurrycane·0
Augmenting Segment customer data with behavioral signals using the Moonsense SDKmoonsense.io·1 pts·hurrycane·0
Moonsense Recorder – Build live prototypes using mobile device sensor datamoonsense.io·2 pts·hurrycane·0
From the Gym to a Jupyter Notebook – Building a Squats Counter App in a Dayurimerhav.medium.com·7 pts·hurrycane·0
AVA: A Finely Labeled Video Dataset for Human Action Understandingresearch.googleblog.com·44 pts·hurrycane·6
Nestlé Targets High-End Coffee by Taking Majority Stake in Blue Bottlemobile.nytimes.com·2 pts·hurrycane·0
Tf-Seq2seq: An Open Source Sequence-To-Sequence Framework in TensorFlowopensource.googleblog.com·3 pts·hurrycane·0
Google Cloud Functions: a serverless environment to connect cloud servicescloudplatform.googleblog.com·1 pts·hurrycane·0