HNHacker News
TopNewBestAskShowJobs

kolinko

5,854 karma · joined September 30, 2009

Entrepreneur from Poland. Founded Szuku.pl, a polish people search engine, and EasyAsk - info-bot company. M.Sc at University of Warsaw. Currently involved in iPhone apps development - www.motivapps.com (first product - a Goal-Setting Workshop: www.motivapps.com/gsw )

workIn-Progress.posterous.com

submissionscomments
kolinko··on OpenAI safety leader quits, warning AI company's culture is 'broken'
Nothing is ever 100% safe, it’s about a right balance of safety to the benefit.

Or, in other words - we have two P(Doom), one for AI being developed, and another for AI being not developed. The latter is not discussed enough imho.

kolinko··on Sites in ChatGPT
They don’t need to, and it would be too risky and too difficult to keep such a conspiracy going.

And also, what’s your story here? That they don’t delete the data they agree to delete, and stash it secretly somewhere? And that they have programs that are mire secretive tan NSA’s inside that secretly finetunes algorithms on such data?

They would expose themselves to lawsuits that would crush their companies, even with such absurd valuations? And possibly risking jail time of the top people? GDPR offences in certain countries are punishable by up to two years in prison. And with a such blatant violation at least in EU the maximum sentences would most likely be given.

kolinko··on Sites in ChatGPT
They can and they do. Unless by “build” you mean guess requirements - then yeah they can’t do it and neither can humans.
kolinko··on Agents don't need memory, they need documentation
What years? Agentic systems, and memory alongside them are roughly only year old.

You’re confusing, I think, memory system with llm finetuning. Completely different concepts.

kolinko··on Sites in ChatGPT
Tech debt when doing a website? If it’s just a CMS, Claude and Codex are more than capable of building and maintaining it.

Also, planning over long scales is not something you should do if you do lean/agile. Premature optimization is a root of all evil ;)

kolinko··on Sites in ChatGPT
Most people host their data on cloud anyway. Once you turn off training there is no difference to storing data with OpenAI/Anthropic vs Google or Meta.
kolinko··on The death of web development education
Slightly disagee.

Personally, I read way less traditional books when it comes to learning stuff. Especially the “utility”/“tutorial”/“handbook”. Top nitch books talking about general principles I’d still read.

Especially for books that need to teach me a subset of a certain discipline - in the past I’d get 4-5 (or at least samples), and try to find the parts that are of interest to me in a style that fits me. Novadays I’d just ask Claude to explain things in a form that I like, with pretty illustrations from Imgen :)

E.g. I don’t think I’ll read the animal books from O’reilly again. But an equivalend of “Thinking in C++” or “Pearls of programming” of a new field? Sure.

kolinko··on AI and the Destruction of the Creative Commons
Well yes, if you can’t articulate what you need then it’s hard for anyone to build it right on the first try.

Otoh I’d say the current models are better at predicting expectations than average programmers. Average programmers don’t know ux or business, LLMs do.

kolinko··on AI and the Destruction of the Creative Commons
Plenty of software on the internet on fully open license (e.g. MIT, copyleft and so on) to train on.

Also, it would be relatively easy to build synthetic datasets for training.

I, for one, don’t mind models being trained on stuff I produced and shared publicly over the last 20 years. I did it for common good, including commercial uses, and this is one of them.

Plenty of people who never produced any open source trying to argue as if if they did.

kolinko··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Why did this get downvotes? :o
kolinko··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
US has 1.4-4.4M programmers and devops, but considering they are more costly than many other professions, it’s quite meaningful to economy. Plus now you have non programmers doing agentic programming to speedup their work - i know nondev project managers, financial advisors, marketers and lawyers rolling their own mini apps now.

Also even if your Gemini is giving you nonprogramming output, underneath the model is most likely generating code for certain tasks.

kolinko··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Such a high cache may be a sign that your agents are using tools inefficiently, thus taking too many turns. Or, like someone else said - you may use tools inefficiently many agents that rediscover stuff etc.

Pi out if the box tries to optimise system prompt size, which is not necessarily good and will cause exactly this effect for all but the most simple tasks.

What you want is to give enough context to the agent to minimize the amount of searching within the codebase etc.

If you want to track cache, what you should do, imho, is to check if you have cache expirations mid sessions (ideally you should not), and if you don’t then lower cache use is actually better - it means that your model doesn’t reread what it just wrote.

kolinko··on The Millennium Problems for Biology
What else?
kolinko··on Measuring the sloppiness of code
IIRC Claude reads PDFs natively, or as images. One page is 3k tokens IIRC, so you can get 100 pages into context and still have plenty of space to reason or work around that. Only for longer and text-specific documents it uses pdftotext, but that's optional.
kolinko··on AI and the Destruction of the Creative Commons
What do you consider "moderately complex software"?
kolinko··on AI and the Destruction of the Creative Commons
What you said doesn't disagree with what the parent said.

LLM can be trained on a code and at the same time reproduce the core ideas. That's what LLMs do after all - they convert the training data into their own internal models and representations, and then reproduce the ideas.

Sure, some things/patterns, that were repeated multiple times, LLMs will tend to repeat verbatim as well, but that's not that big of a problem.

As a person who invented a few algorithms on my own I absolutely love LLMs and I don't mind them being trained on my work, but yeah - I've been way less likely to publish open source over the last year. In the past, if some of my stuff got traction, the credit was close to automatic (early adopters credited or at least knew where they got it from). Nowadays, LLMs will train on these ideas, rewrite them, and give no credit.

Still, I prefer this to having no LLMs at all.

> but stick to the code samples from the books. > I'm certain it'll be able to change the color of a CSS button, right?

A good enough LLM will just decompile a browser, figure out CSS spec from it, and yes - figure out how to change the color of a CSS button from first principles. There is no point to do this with CSS, but with other things it's now easier to just dig through sorces or direct bytecode than to bother checking docs.

kolinko··on I think you should almost never use AI to write
> you should not trust it's "opinions" and "knowledge"

How do you recognize someone as having true opinions and knowledge from someone having just appearance of them?

kolinko··on San Francisco Onion Futures Company
Love it, reminds me of early days of crypto.

How do you decide upon a price?

kolinko··on Measuring the sloppiness of code
If you have more test data I can try this with my harness - you can email me at kolinko@gmail.com btw :)

Thanks for the benchmark, I was seriously looking forward to it! Would you consider such even results good for a one-shot? I wonder how well my harness performs :)

kolinko··on google.com/goto: Google's anti-scraping update
Bombed country? South Korea? Would be Vietnam but US lost there.

Nonbombed countries - a ton, Ukraine for one, but the whole post-communist block.

kolinko··on google.com/goto: Google's anti-scraping update
From where in the rest of the world are you? Did even you talk to anyone living in countries neighboring Russia?

Sure not everyone benefitted from US but a ton of people/countries did and still do.

US never felt a need to build walls to prevent their citizens/allies from leaving. Russia and the others very much so.

With US the track record may be mixed, but a ton of countries from Europe and Asia benefitted tremendously. Russia otoh doesn’t have a concept of win-win, they tend to exploit even their closest allies, which anyone living in Baltics/East-Central Europe can tell you.

kolinko··on Measuring the sloppiness of code
Here’s the oneshot, not sure if it’s slop or not though :)

https://kolinko.eu/pdf-reading-order/

But I wonder about your opinion.

kolinko··on google.com/goto: Google's anti-scraping update
Thanks :)

Another cool book about back and forth between monopolies and decentralisation in our space: https://www.amazon.com/Master-Switch-Rise-Information-Empire...

It also has an audiobook.

kolinko··on google.com/goto: Google's anti-scraping update
As a general rule, if another country has an open war with your neighbor and openly says that you may be next, don’t put in private data into the sites controlled by that country.
kolinko··on google.com/goto: Google's anti-scraping update
It has steep costs, but with monopolists they have other ways of extracting value from the inventions.

Awesome book about the history if Bell Labs - virtually all semiconductor tech we use today was created there (transistors, ics, solar, lasers, fiber optics, telecom satellites…), and they had to license it to be allowed to maintain their monopoly status.

https://www.amazon.pl/Idea-Factory-Great-American-Innovation...

kolinko··on google.com/goto: Google's anti-scraping update
I think it’s the other way around - why should you care about a monopolist’s interests?
kolinko··on google.com/goto: Google's anti-scraping update
Becuse they are a monopolist and we can require certain things from monopolists.

Similarly, Bell Labs was kind of required to release transistor for anyone to license - it was a part of social contract that they were allowed to maintain their monopoly in exchange for releasing certain parts of technology.

Alternatively, they could be split up and their indexing division made an independent company selling to anyone on a free market.

kolinko··on google.com/goto: Google's anti-scraping update
I’d expect everything that you type into Yandex that can be of use to the Russia will be used by Russia - they will nit care about hurting you.

Just like Russia doesn’t care about online criminals being located within their borders - as long as they target outside.

While this may be true to some extent of all the countries, Russia is the one to be clearly against Estonia, Baltics and eastern europeans.

kolinko··on google.com/goto: Google's anti-scraping update
A nice story, but I don’t think there really was a time where closed source models’ development stalled really.
kolinko··on Measuring the sloppiness of code
Yeah that’s why heuristics should work on the lowest possible layer, not on pdftotext. If you use pdftotext you’re stripping positional data and other stuff.

Do you use a public set of documents? I bet I could almost oneshot this with my harness :p

Page 1 of 34Next →