HNHacker News
TopNewBestAskShowJobs

gfourfour

96 karma · joined April 29, 2024

submissionscomments
gfourfour··on Vector indexing all of Wikipedia on a laptop
Yes. I split the text into sentence and append sentences to a chunk until the max context window is reached. The context window size is dynamic for each article so that each chunk is roughly the same size. Then I just do a mean pool of the chunks for each article.
gfourfour··on Vector indexing all of Wikipedia on a laptop
The end result is a recommendation algorithm, so basically just vector similarity search (with a bunch of other logic too ofc). The quality is great, and if anything a little bit of underfitting is desirable to avoid the “we see you bought a toilet seat, here’s 50 other toilet seats you might like” effect.
gfourfour··on Vector indexing all of Wikipedia on a laptop
I will probably make a post when I launch my app! For now I’m trying to figure out how I can host the whole system for cheap because I don’t anticipate generating much revenue
gfourfour··on Vector indexing all of Wikipedia on a laptop
Like 8 gb roughly
gfourfour··on Vector indexing all of Wikipedia on a laptop
Ah didn’t realize it was every language. Yes I’m using a light weight open model - but also my use case doesn’t require anything super heavy weight. Wikipedia articles are very feature-dense and differentiable from one another. It doesn’t require a massive feature vector to create meaningful embeddings.
gfourfour··on Vector indexing all of Wikipedia on a laptop
Nothing too crazy, just downloading a dump, splitting it into manageable batch sizes, and using a lightweight embedding model to vectorize each article. Using the best GPU available on colab it takes maybe 8 hours if I remember correctly? Vectors can be saved as NPY files and loaded into something like FAISS for fast querying.
gfourfour··on Vector indexing all of Wikipedia on a laptop
Yes
gfourfour··on Vector indexing all of Wikipedia on a laptop
Maybe I’m missing something but I’ve created vector embeddings for all of English Wikipedia about a dozen times and it costs maybe $10 of compute on Colab, not $5000
gfourfour··on Humane Explores Potential Sale
Easily the worst demo video I’ve ever seen.

They spoke about the battery life before they even said what the product was supposed to do. The founders sounded like they were being held at gunpoint.

Plus, a $700 price is obscene and a $24/mo subscription is usurious.

That being said, it’s actually a beautiful product. I think maybe given another year it could’ve turned into something workable. I’m assuming they had investors in their ear saying “the first one is never good, but it will fund your next one.”

gfourfour··on Computer-science majors graduate into a world of fewer opportunities
Once you’re in the door somewhere the cream tends to rise, regardless of your background.
gfourfour··on Computer-science majors graduate into a world of fewer opportunities
If you’re interested in doing your own thing send me an email at mhmthrowaway23@gmail.com

I have a BA in philosophy and a MS in CS w/ a focus in computer vision. Similar boat as you. No 13 years of exp in a prior industry though.

I stopped trying to get a job a few months ago and have been plugging away at my own projects. I’ve been (equity compensated) working for a startup for the last month or so in a non-technical role and it’s been really eye opening as to what it takes to get something off the ground. I have a lot of interest in pursuing that route now and if you do too it could be good to connect.

gfourfour··on Computer-science majors graduate into a world of fewer opportunities
Yes, it’s the current market. I’m in the same boat.
gfourfour··on Computer-science majors graduate into a world of fewer opportunities
The market now is terrible but your math degree isn’t holding you back
gfourfour··on Computer-science majors graduate into a world of fewer opportunities
Any kind you want. Being smart matters 100000x more than your major.
gfourfour··on OpenAI created a team to control 'superintelligent' AI – then let it wither
> This paper provides an overview of the main sources of catastrophic AI risks, which we organize into four categories: malicious use, in which individuals or groups intentionally use AIs to cause harm; AI race, in which competitive environments compel actors to deploy unsafe AIs or cede control to AIs; organizational risks, highlighting how human factors and complex systems can increase the chances of catastrophic accidents; and rogue AIs, describing the inherent difficulty in controlling agents far more intelligent than humans.

The first three risks are completely reasonable and people should be thinking about them. No, ChatGPT should not be diagnosing patients and giving them medicine. Yes, we should be vigilant to a flood of disinformation and revenge porn made possible by AI generated content.

But when people talk about “AI safety” in this context, it’s usually in reference to the fourth category, planning for a superintelligent malicious AI that evades detection, self-replicates, etc. That’s pure science fiction at that point, and it’s not a reason to slow down development of LLMs, which yes are basically glorified chatbots and will not lead to “AGI” in this threatening sense.

If I recall correctly, when steam engines started being able to go 40-50 MPH, there were people who were concerned that human beings would not be able to survive travel at such speeds because we never had experienced them. This wasn’t completely irrational, I suppose, as there are speed-induced G forces that are fatal, and they had no way of knowing the threshold back then. But once it was clear that steam locomotives weren’t in any danger of putting us over that threshold, incessant worry about locomotive-induced speeds death was kooky. “Locomotive safety” involving derailment mitigation, track crossing markings, etc. - still legitimate. But if “locomotive safety” was associated with people making claims like “we’re headed for a mass casualty event when the first locomotive hits 60 mph,” then “locomotive safety” would be marginalized.

It doesn’t help that the public faces of “AI safety” include autodidactic pseudointellectuals, clearly mentally unwell people, and philosophers too deep in their own “taken to its logical conclusion…” thought experiments.

gfourfour··on OpenAI created a team to control 'superintelligent' AI – then let it wither
It’s no more thought-terminating than “intelligence,” which is extremely loaded and causes people to make assumptions about these models that work backwards from the “intelligent” label rather than forwards from the tech itself
gfourfour··on Two former MIT students charged with stealing $25M of crypto in 12 seconds
All in the game
gfourfour··on Romance author gets locked out of Google Docs for "inappropriate" content
It’s literally about hockey. Big market for women to read about rich athletic white men.
gfourfour··on Ilya Sutskever to leave OpenAI
This entire saga is really an example of the absurdity of non-profits and philanthropy in general.

The only difference between nonprofit and for-profit entities is that nonprofits divert their profits to a nebulous “cause”, with the investors receiving nothing, while for-profits can distribute profits to their funders.

Other than that, they are free to operate identically.

Generally, entities subject to competitive pressures and with incentives for performance are much better at “benefitting humanity.” Therefore, non-profit status really only makes sense when, one, a profitable enterprise oriented around the intended result isn’t viable (e.g., conservation) or two, there’s a stakeholder that we’ve decided ought to be sheltered from the dynamics of private enterprise, e.g, university students or neutral public broadcasters.

But even in these cases, the non-profit entities basically behave like profit-oriented companies, because their goal is still profitability, just without a return to investors.

OpenAI as a nonprofit would behave the exact same way. There’s no law that the models would have to be open. They’d still be making closed models, charging users, and paying massive salaries. Literally the only difference is that they wouldn’t be able to return money to their investors, and therefore have a much harder time attracting investors, and therefore be less equipped to accomplishing their goal of developing powerful AI.

The irony is that nonprofits are usually only good for things that make for shitty businesses, and things that make shitty businesses usually aren’t that beneficial to humanity. As soon as something becomes really good at what it does, for-profit status makes sense.

What this means, imo, is that most philanthropy dollars are wasted and we would be much better off if they were invested instead. The irony is that this is the point of much philanthropic giving - it ends up being a game of how much money you can burn on nothing, a crass status symbol.

gfourfour··on AI Ruined Quora
To the contrary, the more I use it the more clearly it is just this.

That’s not to disparage the value of a juiced up autocomplete though

gfourfour··on Ask HN: How do you develop and maintain a good note-taking habit?
This was exactly my mindset in high school and to a lesser extent college. I’ve adapted a similar mindset to digital tools with Neovim.

I think you’re right I need to just get back to that system, anything digital is too ethereal to really stick to.

As for pens I’m a hi-tec-C guy btw

gfourfour··on Ask HN: How do you develop and maintain a good note-taking habit?
I agree to an extent but I’ve realized that the benefit of retention is probably outweighed by the downside of slowing information acquisition. I’m realizing that note taking doesn’t help retention as much as I thought it did and not taking notes doesn’t hurt retention as much as a I thought. That’s why I’d like to find more of a middle ground.
gfourfour··on Whistleblower Josh Dean of Boeing supplier Spirit AeroSystems has died
You’re right, that is pedantic
gfourfour··on Investors won't give you the real reason they are passing on your startup
There’s a lot of cargo-cult mimicry of Steve Jobs and Elon Musk that happens in terms of behavior.

But those guys were product people and to the extent that their low-empathy personalities were beneficial it’s in the fact that they cared about creating things far more than anything else, including other people or money. The phrase “artistic temperament” long predates them for a reason.

But it’s really absurd and counterproductive to try to mimic that personality if what you’re offering is anything other than being the corporate version of an auteur, and if that’s what you were you wouldn’t be mimicking anyone, and you better produce some stunningly consistent results. Lots of these guys think they’re being uncompromising about “their vision” when it’s really just “their vision” of their spot on the Forbes list.

New York is better in part, I would guess, because there’s no false premise that the investor is doing anything other than acting as a conduit for capital. There’s egos too but no one is trying to live up to the standard of “changing the world” or “innovating” or being “the next ___.”

It probably also helps that Warren Buffet is the most successful traditional investor of all time and a model of good behavior.

gfourfour··on Ask HN: Nineteen Year Old Needing Advice
It’s pretty normal to have some narcissistic inclinations as a 19 year old, but it’s important to recognize that it’s fundamentally from a place of insecurity and excessive worry about how others perceive you. A good therapist could be helpful.
gfourfour··on They thought they were joining an accelerator – instead they lost their startups
This is kind of a confusing perspective considering VCs are giving you money to pay for the operations of your business. Money is money no matter how stupid the giver is, their money won’t leech value from your company itself.

As for founders ending up with nothing, in those cases their investors ended up with much much less than they were hoping to too. Plus there’s plenty of other cases where founders get rich off a worthless company because of the beneficence of VCs.

gfourfour··on The U.S. economy's big problem? People forgot what 'normal' looks like (2023)
Yeah, it’s not easy to get a job. At all. Maybe I’m insulated in a white-collar bubble, but it seems like there’s a pervasive sense of ‘quiet desperation’ across all industries.
gfourfour··on I Witnessed the Future of AI, and It's a Broken Toy
We need more fun software
gfourfour··on Cloudflare CEO sues over free-roaming fidos at his ski resort paradise
It sounds like they’re walking the dogs on a basically public nature trail that runs through Prince’s property. In which case, in the abstract, I am fine with dogs being off leash.

This is one of those stories where it’s impossible to take a side without accurate context. It could be that the dogs are actually menacing, it could be that Prince isn’t happy about an easement on his land and is using this as a pretext to harass people who rightfully use it to walk.

← PreviousPage 2 of 2