7,435 karma · joined July 22, 2012
It's perhaps the best example I have seen of model drift driven by just small, seemingly unimportant changes to the prompt.
When I checked this a year or so ago, I might have gotten the impression that it was cheaper. Now, it costs the same as what Perplexity charges for search-grounded queries, which is the same as Google charges for Gemini queries with search.
So basically, one player sets a price, and everyone is anchored on that as the pricing for the entire category? I'm just genuinely interested in why every offering in this space is priced like this.
It seems a bit misaligned with how pure LLM queries are priced.
I have a product that would benefit from search grounding, but this pricing wouldn't work with my volume of queries.
I remember reading somewhere that heart transplant recipients have random memory flashes that are not their memories, and sometimes they develop new personality traits.
I'm not saying that I'm a unicorn and that my idea-to-code-to-design execution is flawless, but I certainly believe that in this situation, if I hadn't done it this way, I wouldn't have done it all. However, doing this would be wasteful or dumb in almost every other situation that requires my design output.
People pay for Figma precisely because it's a middle ground. It was a middle ground before, and it will continue to be unless something fundamental changes.
Yes, Apple generates LOTS of revenue overall, but that doesn't justify bleeding cash on a business line that hasn't produced material returns and has no significant positive trajectory in sight.
It's clear that Apple saw this as their Prime Video bet on their services strategy, but that hasn't worked out. Just look at AppleTV+ market share. It's hilariously miniscule.
Despite that, I know people who have ruled out owning a Tesla because they believe the brand mirrors Elon Musk's public persona. They flat-out reject any Tesla product because the brand's visible face is someone they believe doesn't represent their values.
I'm unsure he understood the implications of becoming such a polarizing figure. It was totally unnecessary, yet that was his choice.
I don’t want to write tests for everything. I just want to write the ones that matter.
Make your own conclusions.
Or is something like 1Password truly secure at its core, even if an attacker penetrates some layers of access?
Is something like 1Password truly secure at its core, even if an attacker penetrates some layers of access?
Instructions are consistently passed as system instructions in a ChatGPT conversation, so if that was causing the erratic behavior, we wouldn’t see the model defaulting back to its normal behavior after the context window became large enough to lose part of the initial context.
Every time a newer model is released, we will go through the same cycle of figuring out their emergent intelligent properties over and over.
But I’m not sure that approach will make evident that we are into AGI territory.
We really are going to need new kind of evaluations because it’s evident that passing the bar or whatever isn’t really give you a proxy for intelligence, let alone sentience.
But I do really like that functionality of attaching context to the query! Love that. I replied to the founder of Sourcegraph in this thread and I think that should answer your question as well.
I'm excited about your approach and even more about the fact that you made it open source! Thanks for that. If this sticks for me I can see myself contributing to it.
Let's say I want to build a complex React component that does a lot of stuff under the hood. Perhaps it needs to handle multiple inputs, it needs to show certain children UI conditionally based on the state and it also needs to push the state to a global redux store.
Perhaps the component also depends on some utils and other root components that pass props to it.
None of these solutions seem to acknowledge that what would improve my productivity is to have a proper way to model the problem for the AI the same way I'm understanding it. Allowing me to build the blast radius of the problem, instead of expecting the AI to infer it.
In the case of the React Component what I want is:
1) bootstrap the component
2) modify it to address requirements that emerge as I explore the needs.
3) look at dependencies such as schema files or root components and suggest modifications that align with the desired functionality or allow me to point out at required changes in dependencies.
When I say that these solutions are suboptimal, what I mean is that there's no straight-forward way for me to engage in these code generation tasks in a way that doesn't feel fragmented.
What I literally want is go to a file, tell the AI to consume all the context of that file and its dependencies and then modify it to add features, fix bugs, improve code or extend functionality. I then want to have the ability to accept those changes as a reviewer or suggest changes. And I want to do without forking away from my current context.
Nobody has gotten this experience right because what most players in this space are doing is building these restrictive form factors or atomized features like text brushes, global chats or in-line prompt to code generation.
I appreciate having these, but ultimately what I want is to have an AI system that accepts a context (a file or set of files) and a prompt and then generates the requested code modifications.
None of these extension allow me to do this effectively and if they do then it seems that ability is being diminished by a poor user-facing abstraction.
In my opinion, this is a design problem.
I have tried almost every co-pilot solution, including GitHub Co-Pilot with Chat and Labs, Cody, and a few random extensions.
And for some reason, I still default to using ChatGPT, even with the massive drawback of copying and pasting.
I haven’t seen any breakthroughs with these developer experiences. Every still feels suboptimal.
I'm not sure it's entirely clear what the Cybertruck means for Tesla or the car industry, but it's obviously not a minor event in this landscape.
Especially after so much speculation on whether or not this truck would ever become a production car.
I have always stayed away from that region because it seems significantly less reliable than other regions.
On the other hand, I guess VC is just that. To follow trends, predict trajectories and attempt to make one win out of thousands of investmens. But a16z trying to become a thought leader in AI after all the crypto BS they elevated is very off-putting.
Times are different now. If I was a founder in AI, I would probably be wary of firms that went so hard on crypto. It seems their investment thesis is to monopolize attention around these trends instead of seeking real alignment.
https://engineering.fb.com/2017/03/29/data-infrastructure/fa...
Then I changed the prompt slightly, and it answered that it supports 512 tokens contradicting its previous answer.
That's like early GPT-3.0 level performance, including a good dose of hallucinations.
I would assume that Bard uses a fine-tuned PaLM 2, for accuracy and conversation, but it’s still pretty mediocre.
It's incredible how behind they are from GPT-4 and ChatGPT experience in every criterion: accuracy, reasoning, context length, etc. Bard doesn't even have character streaming.
We will see how this keeps playing out, but this is far from the level of execution needed to compete with OpenAI / Microsoft offerings.
That's where I stopped reading. GitLab’s Merge Requests are on par or even more robust than GitHub’s Pull Requests.
This person clearly hasn't used GitLab; otherwise, they wouldn’t say this.
Blank statements like this one undermine their core argument.