HNHacker News
TopNewBestAskShowJobs

whoisjuan

7,435 karma · joined July 22, 2012

Opinions and thoughts from a random internet stranger. Take with a grain of salt.
submissionscomments
whoisjuan··on AI World Clocks
The time given to the model. So the difference between two generations is just somethng trivially different like: "12:35" vs 12:36"
whoisjuan··on AI World Clocks
It's actually quite fascinating if you watch it for 5 minutes. Some models are overall bad, but others nail it in one minute and butcher it in the next.

It's perhaps the best example I have seen of model drift driven by just small, seemingly unimportant changes to the prompt.

whoisjuan··on Launch HN: Exa (YC S21) – The web as a database
Did you guys change the pricing of Exa?

When I checked this a year or so ago, I might have gotten the impression that it was cheaper. Now, it costs the same as what Perplexity charges for search-grounded queries, which is the same as Google charges for Gemini queries with search.

So basically, one player sets a price, and everyone is anchored on that as the pricing for the entire category? I'm just genuinely interested in why every offering in this space is priced like this.

It seems a bit misaligned with how pure LLM queries are priced.

I have a product that would benefit from search grounding, but this pricing wouldn't work with my volume of queries.

whoisjuan··on Memories are not only in the brain, human cell study finds
https://pubmed.ncbi.nlm.nih.gov/38694651/
whoisjuan··on Memories are not only in the brain, human cell study finds
This is wild, but many studies have reached the same conclusion.

I remember reading somewhere that heart transplant recipients have random memory flashes that are not their memories, and sometimes they develop new personality traits.

whoisjuan··on DoNotPay has to pay $193K for falsely touting untested AI lawyer, FTC says
This is a very sneaky ethically gray company. Their app is not only of terrible quality but also full of dark patterns. I'm convinced that any revenue they make comes from people who can't figure out how to cancel. Stay away from it.
whoisjuan··on Will Figma become an awkward middle ground?
Ohh nice catch. Will do. Thanks
whoisjuan··on Will Figma become an awkward middle ground?
I'm a designer. I built brainglue.ai without Figma, a design system, or a UI library. I just went directly to code (react+tailwind) and let a style organically emerge.

I'm not saying that I'm a unicorn and that my idea-to-code-to-design execution is flawless, but I certainly believe that in this situation, if I hadn't done it this way, I wouldn't have done it all. However, doing this would be wasteful or dumb in almost every other situation that requires my design output.

People pay for Figma precisely because it's a middle ground. It was a middle ground before, and it will continue to be unless something fundamental changes.

whoisjuan··on Apple tries to rein in Hollywood spending after years of losses
AppleTV+ is a tiny business. It's nowhere near of generating enough revenue to cover a $20B hole in content production costs.

Yes, Apple generates LOTS of revenue overall, but that doesn't justify bleeding cash on a business line that hasn't produced material returns and has no significant positive trajectory in sight.

It's clear that Apple saw this as their Prime Video bet on their services strategy, but that hasn't worked out. Just look at AppleTV+ market share. It's hilariously miniscule.

whoisjuan··on Walls are starting to close in for Tesla, let's have a closer look
I own a Model 3, and I honestly think it is the best car I have ever owned.

Despite that, I know people who have ruled out owning a Tesla because they believe the brand mirrors Elon Musk's public persona. They flat-out reject any Tesla product because the brand's visible face is someone they believe doesn't represent their values.

I'm unsure he understood the implications of becoming such a polarizing figure. It was totally unnecessary, yet that was his choice.

whoisjuan··on Meta's new LLM-based test generator
No op, but I don’t think test-driven development resounds with everyone who writes code.

I don’t want to write tests for everything. I just want to write the ones that matter.

whoisjuan··on GPT-4-turbo produces shorter completions when it "thinks" its December vs. May
I think productivity is lower in the winter, so I'm not sure about quality per se, but intuitively it makes sense that anything written in the winter months is less verbose.
whoisjuan··on OpenAI's board has fired Sam Altman
GPTs is basically a ripoff of Poe by Quora. Quora’s CEO is Adam D’ Angelo. Adam D’ Angelo is one of OpenAI’s board members.

Make your own conclusions.

whoisjuan··on 1Password detects "suspicious activity" in its internal Okta account
It was a minor incident, but it does remind me that centralized password managers seem to have an awful amount of concentrated risk.

Or is something like 1Password truly secure at its core, even if an attacker penetrates some layers of access?

whoisjuan··on 1Password discloses security incident linked to Okta breach
It was a minor incident, but it does remind me that centralized password managers seem to have an awful amount of concentrated risk.

Is something like 1Password truly secure at its core, even if an attacker penetrates some layers of access?

whoisjuan··on What happened in this GPT-3 conversation?
Unlikely. You can see that the model returns to normal behavior after it exhausts the context window that causes this.

Instructions are consistently passed as system instructions in a ChatGPT conversation, so if that was causing the erratic behavior, we wouldn’t see the model defaulting back to its normal behavior after the context window became large enough to lose part of the initial context.

whoisjuan··on What happened in this GPT-3 conversation?
One thing I have observed with the rise of generative AI is that the general direction everyone is pushing towards is to make the models behave deterministically when in principle, LLMs are probabilistic.

Every time a newer model is released, we will go through the same cycle of figuring out their emergent intelligent properties over and over.

But I’m not sure that approach will make evident that we are into AGI territory.

We really are going to need new kind of evaluations because it’s evident that passing the bar or whatever isn’t really give you a proxy for intelligence, let alone sentience.

whoisjuan··on Show HN: Continue – Open-source coding autopilot
Hey. I gave it a try to Continue and I think this is going in the right direction at least for me. I guess opinions on how this should work are subjective.

But I do really like that functionality of attaching context to the query! Love that. I replied to the founder of Sourcegraph in this thread and I think that should answer your question as well.

I'm excited about your approach and even more about the fact that you made it open source! Thanks for that. If this sticks for me I can see myself contributing to it.

whoisjuan··on Show HN: Continue – Open-source coding autopilot
I think there are several problems with these co-pilot solutions, but the most obvious ones are the context switching problem and the lack of steer-ability.

Let's say I want to build a complex React component that does a lot of stuff under the hood. Perhaps it needs to handle multiple inputs, it needs to show certain children UI conditionally based on the state and it also needs to push the state to a global redux store.

Perhaps the component also depends on some utils and other root components that pass props to it.

None of these solutions seem to acknowledge that what would improve my productivity is to have a proper way to model the problem for the AI the same way I'm understanding it. Allowing me to build the blast radius of the problem, instead of expecting the AI to infer it.

In the case of the React Component what I want is:

1) bootstrap the component

2) modify it to address requirements that emerge as I explore the needs.

3) look at dependencies such as schema files or root components and suggest modifications that align with the desired functionality or allow me to point out at required changes in dependencies.

When I say that these solutions are suboptimal, what I mean is that there's no straight-forward way for me to engage in these code generation tasks in a way that doesn't feel fragmented.

What I literally want is go to a file, tell the AI to consume all the context of that file and its dependencies and then modify it to add features, fix bugs, improve code or extend functionality. I then want to have the ability to accept those changes as a reviewer or suggest changes. And I want to do without forking away from my current context.

Nobody has gotten this experience right because what most players in this space are doing is building these restrictive form factors or atomized features like text brushes, global chats or in-line prompt to code generation.

I appreciate having these, but ultimately what I want is to have an AI system that accepts a context (a file or set of files) and a prompt and then generates the requested code modifications.

None of these extension allow me to do this effectively and if they do then it seems that ability is being diminished by a poor user-facing abstraction.

In my opinion, this is a design problem.

whoisjuan··on Show HN: Continue – Open-source coding autopilot
Nice! I didn't know about this! Thanks.
whoisjuan··on Show HN: Continue – Open-source coding autopilot
Thanks! I’ll give this a try for sure.
whoisjuan··on Show HN: Continue – Open-source coding autopilot
Why is nobody in this space making good gains on UX?

I have tried almost every co-pilot solution, including GitHub Co-Pilot with Chat and Labs, Cody, and a few random extensions.

And for some reason, I still default to using ChatGPT, even with the massive drawback of copying and pasting.

I haven’t seen any breakthroughs with these developer experiences. Every still feels suboptimal.

whoisjuan··on Ford slashes prices of F-150 Lightning trucks as EV wars heat up
I think Ford might be scared of the cult-like culture that has been created around Tesla and/or Elon Musk.

I'm not sure it's entirely clear what the Cybertruck means for Tesla or the car industry, but it's obviously not a minor event in this landscape.

Especially after so much speculation on whether or not this truck would ever become a production car.

whoisjuan··on Ford slashes prices of F-150 Lightning trucks as EV wars heat up
Knee-jerk reaction to Cybertruck's imminent launch? https://cbsaustin.com/news/local/teslas-austin-based-gigafac...
whoisjuan··on AWS us-east-1 down
Why is it always us-east-1 though?

I have always stayed away from that region because it seems significantly less reliable than other regions.

whoisjuan··on AI will save the world?
Is it me or a16z completely destroyed their reputation as a reputable VC after the crypto frenzy? I just can't seem to read this with a straight face.

On the other hand, I guess VC is just that. To follow trends, predict trajectories and attempt to make one win out of thousands of investmens. But a16z trying to become a thought leader in AI after all the crypto BS they elevated is very off-putting.

Times are different now. If I was a founder in AI, I would probably be wary of firms that went so hard on crypto. It seems their investment thesis is to monopolize attention around these trends instead of seeking real alignment.

whoisjuan··on ChatGPT plugins now support Postgres and Supabase
Does this use Faiss to find similar vectors?

https://engineering.fb.com/2017/03/29/data-infrastructure/fa...

whoisjuan··on PaLM 2 Technical Report [pdf]
That's because simple questions in Bard only generate like 200 tokens per answer. The latency is more noticeable for longer answers.
whoisjuan··on PaLM 2 Technical Report [pdf]
Well, I tried it, and this is how dumb it is. I ask it what's the context length it supports. It said that PaLM 2 supports 1024 tokens and then proceeds to say that 1024 tokens equals 1024 words, which is obviously wrong.

Then I changed the prompt slightly, and it answered that it supports 512 tokens contradicting its previous answer.

That's like early GPT-3.0 level performance, including a good dose of hallucinations.

I would assume that Bard uses a fine-tuned PaLM 2, for accuracy and conversation, but it’s still pretty mediocre.

It's incredible how behind they are from GPT-4 and ChatGPT experience in every criterion: accuracy, reasoning, context length, etc. Bard doesn't even have character streaming.

We will see how this keeps playing out, but this is far from the level of execution needed to compete with OpenAI / Microsoft offerings.

whoisjuan··on Let's make sure GitHub doesn't become the only option
“ You can see similar features in GitHub’s smaller competition Codeberg, GitLab, BitBucket, and Gitea. These competitors don’t offer other, major code collaboration tools, and their Pull Request-like features aren’t just there to help users come from GitHub.“

That's where I stopped reading. GitLab’s Merge Requests are on par or even more robust than GitHub’s Pull Requests.

This person clearly hasn't used GitLab; otherwise, they wouldn’t say this.

Blank statements like this one undermine their core argument.

Page 1 of 34Next →