HNHacker News
TopNewBestAskShowJobs

fogx

14 karma · joined January 23, 2020

submissionscomments
fogx··on Review: Chuwi's $449 Unibook laptop is a funhouse-mirror MacBook Neo
Yea, moving the burden from your wrists to your finger joints is surely a genius idea and the epitome of health
fogx··on Use One Big Server (2022)
don't you think it's highly unlikely that someone will stumble over the power cable in a hosted datacenter like hetzner? and even if, you could just run a provisioned secondary server that jumps in if the first becomes unavailable and still be much cheaper.
fogx··on Is 4chan the perfect Pirate Bay poster child to justify wider UK site-blocking?
yea right. Privacy is a fundamental right in the EU (GDPR, Charter of Fundamental Rights), while the U.S. legal system offers almost no general privacy protection. On top of that, the NSA has a long history of warrantless surveillance and backdoors (Snowden, PRISM), with very limited oversight. In practice, it’s far costlier to push mass privacy infringements in Europe than in the U.S.
fogx··on Web-scraping AI bots cause disruption for scientific databases and journals
esp. for image data libraries, why not provide the images as a dump instead? No need to crawl 3mil images if the download button is right there. Now put the file on a cdn or Google and you're golden
fogx··on Show HN: 1 min workouts for people who sit all day
webapp would have been nice :(

where do you pull the exercises from?

fogx··on Understanding the BM25 full text search algorithm
we use it and are fairly happy. but provider latency is insanely high (500ms+) for embedding models. best to host on-cluster. hybrid quality is good but modification options are extremely limited and the score very obscure for anything but ranking within the set.
fogx··on Google’s AI thinks I left a Gatorade bottle on the moon
yes, it is
fogx··on Launch HN: Haystack (YC S24) – Visualize and edit code on an infinite canvas
seems like a neat idea, but how does this 10x my dev performance?
fogx··on Show HN: Visualizing Chess Games
looks cool, how can I use this information?
fogx··on How are generative AI companies monitoring their systems in production?
Do you have any links to conversational simulation?
fogx··on How long it took different companies to find product-market fit
multiple quotes mentioned tracking early user activity on a very granular level. Any recommendations for such analytics? I know only hotjar
fogx··on Advanced NLP with SpaCy
> For each predicted output token, I want to know exactly which source document(s) were utilized including indices from those documents and relevant statistics.

You don't do this for any other kind of decision or tool, why do you need it for LLMs?

i think in truth you need a source that convinces you (or the regulator) that your choice is acceptable, so that you can pass off the responsibility. If the LLM were to give you an answer backed by relevant (cited) source documents of regulations and a good explanation it would make no difference compared to a human worker doing the same. -> this is already possible

fogx··on Yes, you can measure software developer productivity
SMART goals and KPIs are meant to improve work focus, not work ethos. Those two are not in opposition to each other. IME engineers who don't want to do measurably good work are inexperienced, disgruntled or unmotivated.

Forcing goals onto a disgruntled/unmotivated engineer will end badly. Giving no goals to motivated engineers will also end badly. Giving bad, unrelated or unachievable goals to motivated engineers will end in disgruntled or unmotivated engineers.

You have to get both right.

fogx··on Ask HN: Who is using small OS LLMs in production?
have you considered azure's GPT, or is that not private enough?
fogx··on Hard stuff when building products with LLMs
> especially if those inputs are extremely vague (how on earth should we interpret slow?!)

isn't this exactly the (theoretical) strength of a chatbot - asking follow-up questions to remove uncertainty?

fogx··on How and why I built Japan Dev
Thanks for your input! server-side rendering should fix most of the first issue. You can get all the relevant data pre-rendered and it's pretty easy to do with the right framework. I've used quasar.dev for this in the past. It should also solve the initial load times (to an extent). The chaining can be fixed with lazy loading/code splitting.

The only thing i haven't fixed yet is hiding the cookie banner for bots :/

fogx··on Vector search just got up to 10x faster and vertically scalable
This sounds like a perfect job for jina.ai -> sharding, redudancy, automatic up- and down scaling, security features and most importantly: flexibility to use and switch whatever (vector) database https://jina.ai/
fogx··on How and why I built Japan Dev
Can you explain more about these headaches? So far i'm still convinced of client side SPAs :D
fogx··on The AI Art Apocalypse
image-to-text models for captioning already exist. The most common one is CLIP from openAI. https://openai.com/blog/clip/ Jina AI has an out-of-the-box implementation for it https://clip-as-service.jina.ai/