HNHacker News
TopNewBestAskShowJobs

dchichkov

2,070 karma · joined November 17, 2011

engineer. researcher. founder.

... "What I can not create, I do not understand." ...

... "Know how to solve every problem that has been solved." ...

... "Four no's. Five clues." ...

... "Reality is that which, when you stop believing in it, doesn’t go away." ...

... "A human being is a part of the whole called by us universe, a part limited in time and space." ...

submissionscomments
dchichkov··on OpenAI asks White House for relief from state AI rules
>> In the proposal, OpenAI also said the U.S. needs “a copyright strategy that promotes the freedom to learn” and on “preserving American AI models’ ability to learn from copyrighted material.”

Perhaps also symmetric "freedom to learn" from OpenAI models, with some provisions / naming convention? U.S. labs are limited in this way, while labs in China are not.

dchichkov··on Music labels will regret coming for the Internet Archive, sound historian says
Just in case, here's the list of these labels:

- UMG Recordings, Inc.

- Capitol Records, LLC

- Concord Bicycle Assets, LLC

- CMGI Recorded Music Assets LLC

- Sony Music Entertainment

- Arista Music

Taken from: https://cdn.arstechnica.net/wp-content/uploads/2025/03/UMG-v...

dchichkov··on ARC-AGI without pretraining
0 1 00 01 10 11 000 001 010 011 100 101 110 111 0000 0001 0010 0011 0100 0101 0110 0111 1000 1001 1010 1011 1100 1101 1110

And no, I don't think the knowledge of language is necessary. To give a concrete example, tokens from TinyStories dataset (the dataset size is ~1GB) are known to be sufficient to bootstrap basic language.

dchichkov··on ARC-AGI without pretraining
For long context sizes AGI is not useless without vast knowledge. You could always put a bootstrap sequence into the context (think Arecibo Message), followed by your prompt. A general enough reasoner with enough compute should be able to establish the context and reason about your prompt.
dchichkov··on The Deep Research problem
I agree, they are only starting the data flywheel there. And at the same time making users pay $200/month for it, while the competition is only charging $20/month.

And note, the system is now directly competing with "interns". Once the accuracy is competitive (is it already?) with an average "intern", there'd be fewer reasons to hire paid "interns" (more expensive than $200/month). Which is maybe a good thing? Fewer kids wasting their time/eyes looking at the computer screens?

dchichkov··on [dead]
The approach of "cutting funding and then observing whether anything critical fails or is impacted" only works if outcomes follow a normal distribution.

This is far from the case — many areas are characterized by heavy-tailed loss distributions, where extreme negative consequences could really ruin the day and erase any efficiency gains.

dchichkov··on OpenAI says it has evidence DeepSeek used its model to train competitor
I've suggested that long context should be included into the prompt.

In your particular case the prompt would look something like: <pubmed dump> what are the plants that aren't poisonous to most people?

A general reasoner would recover language and relevant world model from pubmed dump. And then would proceed to reason about it, to perform the task.

It doesn't look like a particularly efficient process.

dchichkov··on OpenAI says it has evidence DeepSeek used its model to train competitor
If you look at the benchmarks of the DeepSeek-V3-Base, it is quite capable, even in 0-shot: https://huggingface.co/deepseek-ai/DeepSeek-V3-Base#base-mod... This is not from scratch. These benchmark numbers are an indication that the base model already had a large number of reasoning/LLM tokens in the pre-training set.

On the other hand, my take on it, the ability to do reasoning in a long context is a general capability. And my guess is that it can be bootstrapped from scratch, without having to do training on all of the internet or having to distill models trained on the internet.

dchichkov··on OpenAI says it has evidence DeepSeek used its model to train competitor
There are examples of learning reasoning from scratch with reinforcement learning.

Emergent tool use from multi-agent interaction is a good example - https://openai.com/index/emergent-tool-use/

dchichkov··on OpenAI says it has evidence DeepSeek used its model to train competitor
It should be possible to learn to reason from scratch. And the ability to reason in a long context seems to be very general.
dchichkov··on DeepSeek releases Janus Pro, a text-to-image generator [pdf]
MMMU is not particularly high. Janus-Pro-7B is 41.0, which is only 14 points better than random/frequent choice. I'm pretty sure, their base DeepSeek 7B LLM will get around 41.0 MMMU without access to images, this is a normal number for a roughly GPT4-level LLM base with no access to images.
dchichkov··on Prime numbers so memorable that people hunt for them
Sorry, but this was ChatGPT/o1 with access to code execution (Python) and it used almost 4 minutes to do reasoning. It had done a few checks with smaller numbers, all of which had failed. And it proceeded to make a wrong conclusion (with high confidence).
dchichkov··on Prime numbers so memorable that people hunt for them
ChatGPT o1: https://chatgpt.com/share/678feedb-0b2c-8001-bd77-4e574502e4...

> Thought about large prime check for 3m 52s: "Despite its interesting pattern of digits, 12,345,678,910,987,654,321 is definitely not prime. It is a large composite number with no small prime factors."

Feels like this Online Encyclopedia of Integer Sequences (OEIS) would be a good candidate for a hallucination benchmark...

dchichkov··on Ask HN: How do you prevent the impact of social media on your children?
I understand that it is mostly regulated at the state level. I'm not sure about other states, but The Computer Science Standards for California Public Schools (Kindergarten through Grade Twelve) also tend to be followed by private schools. So they can claim their programs meet state requirements.

This brings computers into the classroom, and once they’re available, it is a slippery slope. It is easier for teachers to have students use semi-gamified "educational" apps rather than engage themselves.

Example for K-2 - https://www.cde.ca.gov/be/st/ss/documents/csstandards.pdf:

  K-2.CS.1 Select and operate computing devices that perform a variety of tasks accurately and quickly based on user needs and preferences.

  K-2.CS.2 Explain the functions of common hardware and software components of computing systems.

  K-2.CS.3 Describe basic hardware and software problems using accurate terminology.

  K-2.NI.4 Model and describe how people connect to other people, places, information and ideas through a network.

  ...

  K–2 K-2.AP.12 Create programs with sequences of commands and simple loops, to express ideas or address a problem

  K-2.IC.20 Describe approaches and rationales for keeping login information private, and for logging off of devices appropriately
dchichkov··on Ask HN: How do you prevent the impact of social media on your children?
Another Gorilla is the schools, teachers and state-approved recommendations, that extend their reach even into private schools.

Imagine my frustration one day, when I've discovered that my kindergartner has full access to a brand-new, shiny iPad during class. Despite complaints from parents, the teacher refused to reduce iPad usage (or even activate Screen Distance and Screen Time controls on the iPad, or share usage statistics).

The only thing that I've learned, this is all in line with California’s state-approved computer literacy recommendations.

dchichkov··on FTC bans hidden junk fees in hotel, event ticket prices
At the expense of other people's time.
dchichkov··on FTC bans hidden junk fees in hotel, event ticket prices
Whoever invented this is evil and is destroying happiness.
dchichkov··on FTC bans hidden junk fees in hotel, event ticket prices
Safeway, Walgreens.
dchichkov··on FTC Announces Rule Banning Junk Ticket and Hotel Fees
I wish that "Online Coupon Price Tags" in stores would also be banned. I'm talking about these yellow price tags that show lower than "Club" prices, which are only valid if you collect a coupon online.

Like FTC, I estimate that banning these would save U.S. consumers millions of hours they currently spend searching and clicking on pointless coupons on their phones before making purchases. It would also increase happiness, as it's extremely annoying to pay $20 extra, knowing that a lower price is available if only you spent ten minutes struggling with a store's website on your phone.

Whoever invented this is evil and is destroying happiness.

dchichkov··on FTC bans hidden junk fees in hotel, event ticket prices
I wish that "Online Coupon Price Tags" in stores would also be banned. I'm talking about these yellow price tags that show lower than "Club" prices, which are only valid if you collect a coupon online.

Like FTC, I estimate that banning these would save U.S. consumers millions of hours they currently spend searching and clicking on pointless coupons on their phones before making purchases. It would also increase happiness, as it's extremely annoying to pay $20 extra, knowing that a lower price is available if only you spent ten minutes struggling with a store's website on your phone.

Whoever invented this is evil and is destroying happiness.

dchichkov··on Phased Array Microphone (2023)
Yeah, I've also had difficulty finding something with enough I2S. It was a while back and I've used Sprocket carrier for Jetson TX2 - it had 6 lanes, so up to 96. It was for a SODAR application, so the sampling frequency was not that critical and to me it felt like the perfect trick to make an array with off-the-shelf hardware. So I was just curious, if this was something you've considered.

For something indoors, yes, I can see how low sampling frequency gets very limiting. And 192 microphones, that's really pushing it. Love it.

The $2/mic vs $0.5/mic argument is a fun one. You've obviously poured enormous amount of engineering in there, involving PCB design, FPGA and network programming, writing custom CUDA kernels, signal processing, PyTorch, the list goes on. And you've had 4090 plugged in your PC in 2023. Classic hobbit in a mithril vest ;)

dchichkov··on Phased Array Microphone (2023)
I'm curious, why haven't you used TDM I2S microphones for your array and used PDM?

I understand that ICS-52000 is a relatively low cost ($2/100pcs) and there are even breakout boards available with 4 microphones, which can be chained to 8 or 16, like https://www.cdiweb.com/datasheets/notwired/ds-nw-aud-ics5200...

Then you can take Jetson (or any I2S capable hardware with DSP or GPU on it) and chain 16 microphones per I2S port. It would seem a lot easier to assemble and probgam, if comared to FPGA setup.

dchichkov··on Google AI chatbot responds with a threatening message: "Human Please die."
It was a long chat - https://g.co/gemini/share/6d141b742a13 and then the last question contained text that was completely broken. It's not surprising to have a failure case under such conditions.

And it is reasonable to have failure cases. But systems should fail gracefully. This wasn't a graceful failure.

dchichkov··on AI isn't unleashing imaginations, it's outsourcing them
There's nothing magical about AI, but a forward pass through a transformer is a rather large stochastic computation. During this computation novel results are sometimes produced.

If you'd like an insight of what roughly can happen inside this computation, a short story from Karpathy is not the worst read - http://karpathy.github.io/2021/03/27/forward-pass/

dchichkov··on Decree 770
It is an interesting piece of history. It was mentioned in a middle of a rather long book: "Behave: The Biology of Humans at Our Best and Worst", and I was surprised that it is not on the surface and I didn't know about it before. Considering how relevant it is to the current happenings in the US.
dchichkov··on Tesla's Cybertruck is outselling almost every other EV in the US
Is there some tax data about the amount of Clean Vehicle/CA and Federal EV credits issued for Cybertruck?

The income cap on getting the clean vehicle rebates is $135k ($200k joint filers). And I'm not sure about the federal rebates. Tesla doesn't offer 0% financing, current Cybertruck APR deal is reported to be 5.29% for up to 72 months. So I don't see how someone with the income under the rebate cutoff can afford that $100k car or the financing option. The delta between the number of rebates (Federal EV vs Clean Vehicle/CA) may allow to estimate, how many of these are corporate (pre-income tax + rebate?) purchases.

And these "Cox Automotive estimates", are these reliable numbers that had been confirmed by Tesla earnings, or it is a "best guess by influencers" type of information?

dchichkov··on US probes Tesla's Full Self-Driving software after fatal crash
> I'm grateful to be getting a car from another manufacturer this year.

I'm curious, what is the alternative that you are considering? I've been delaying an upgrade to electric for some time. And now, a car manufacturer that is contributing to the making of another Jan 6th, 2021 is not an option, in my opinion.

dchichkov··on NASA freezes Starliner missions
As long as there's competition, it is fine. Boeing fits at least that role easily. Plus, they've built the vehicle with no drama and without purchasing Twitter in the middle. This is worth something.

We see similar situation in automotive. Other companies do allow to keep Tesla in check, so there's less opportunity to force "Cybertrucks" onto the market as the only option.

dchichkov··on NASA freezes Starliner missions
It is unhealthy to not have competition to SpaceX.
dchichkov··on The Open Source AI Definition RC1 Is Available for Comments
I agree, Open Weights are Open "Binary", not Open Source.

It's like taking an executable (.so module, firmware blob) and releasing it under permissive license, so anyone could disassemble, modify and hack it. And then disclosing what programming languages were used and pointing at a few libraries. And then saying that no, actual source code is not going to be released.

Page 1 of 28Next →