GitHub cuts AI deals with Google, Anthropic
bloomberg.com
bloomberg.com
I find that ai can help significantly with doing plumbing, but it has no problems with connecting the pipes wrong. I need to double and triple check the updated code - or fix the resulting errors when I don’t do that. So: boilerplate and outer app layers, yes; architecture and core libraries, no.
Curious, is that a property of all ai assisted tools for now? Or would copilot, perhaps with its new models, offer a different experience?
My theory is the willingness to baby sit and the modality. I'm perfectly fine telling the tool I use its errors and working side by side with it like it was another person. At the end of the day it can belt out lines of code faster than I, or any human, can and I can review code very quickly so the overall productivity boost has been great.
It does fundamentally alter my workflow. I'm very hands off keyboard when I'm working with AI in a way that is much more like working with someone or coaching someone to make something instead of doing the making myself. Which I'm fine with but recognize many developers aren't.
I use AI autocomplete 0% of the time as I found that workflow was not as effective as me just writing code, but most of my most successful work using AI is a chat dialogue where I'm letting it build large swaths of the project a file or parts of a file at a time, with me reviewing and coaching.
The apparent speed up is mostly a deception. It definitely helps with rough outlines and approaches. But, the faster you go, the less you will notice the fine details, and the more assumptions you will accumulate before realizing the fundamental error.
I'd rather find out I was wrong within the same day. I'd probably have written some unit tests and played around with that function a lot more if I had handcrafted it.
If it wants to complete what I wanted to type anyway, or something extremely similar, I just press tab, otherwise I type my own code.
I'd say about 70% of individual lines are obvious enough if you have the surrounding context that this works pretty well in practice. This number is somewhat lower in normal code and higher in unit tests.
Another use case is writing one-off scripts that aren't connected to any codebase in particular. If you're doing a lot of work with data, this comes in very handy.
Something like "here's the header of a CSV file", pass each row through model x, only pass these three fields, the model will give you annotations, put these back in the csv and save, show progress, save every n rows in case of crashes, when the output file exists, skip already processed rows."
I'm not (yet) convinced by AI writing entire features, I tried that a few times and it was very inconsistent with the surrounding codebase. Managing which parts of the codebase to put in its context is definitely an art though.
It's worth keeping in mind that this is the worst AI we'll ever have, so this will probably get better soon.
I'd highly recommend reading through Aider's docs[0], because I think it's relevant for any AI tool you use. A lot of people harp on prompting, and while a good prompt is important I often see developers making other mistakes like not providing context that's good, correct, or even too much[1].
When I find models are going on the wrong path with something, or "connecting the pipes wrong", I often add code comments that provide additional clarity. Not only does this help future me/devs, but the more I steer AI towards correct results, the fewer problems models seem to have going forward.
Everybody seems to be having wildly different experiences using AI for coding assistance, but I've personally found it to be a big productivity boost.
[0] https://aider.chat/docs/usage/tips.html
[1] https://aider.chat/docs/troubleshooting/edit-errors.html#red...
But working on our actual codebase with copilot in the IDE (Rider, in my case) is a net negative. It usually does OK when it's suggesting the completion of a single line, but when it decides to generate a whole block it invariably misunderstands the point of the code. I could imagine that getting better if I wrote more descriptive method names or comments, but the killer for me is that it just makes up methods and method signatures, even for objects that are part of publicly documented frameworks/APIs.
it can be used if you want the reliability of a random forum poster. which... sure. knock yourself out. sometimes there's gems in that dirt.
I'm getting _very_ bearish on using LLMs for things that aren't pattern recognition.
Really not worth it for me at this point.
I must add that I found Anthropic's Claude quite useful (Sonnet 3.5) when I had to work on a legacy code base using Adobe ColdFusion (*vomit). I knew nothing of Coldfusion or the awful code base, and it helped me figure out a lot of things about Coldfusion without having to spend too much energy on learning and generating code for a framework I will never again use. I still had to make some updates to the code, but it was less cognitive effort that having to read docs and spend time Googling.
How can this be possible if you literally admit its tab completion is mindblowing?
Isn't really good tab completion good enough for at least a 5% producitvity boost? 10%? 20%?
Select line of code, prompt it to refactor, verify they are good, accept the changes
Just sanity checking that the output and “piping” is correct.
My productivity (in frontend work at least) is significantly higher than before.
What I notice is that line-completion is quite good if you read the produced code in prosa. But most of the times the internals are different, and so the completed line is still useless.
E.g. assume an Enum Completion with values InlineCompletion, FullLine, NoCompletion.
If i write
> if (currentMode == Compl
It will happily suggest
> if (currentMode == Completion.FullLineCompletion) {
while not realizing that this enum values does not exist.
SpaceX's advancements are impressive, from rocket blow up to successfully catching the Starship booster.
Who knows what AI will be capable of in 5-10 years? Perhaps it will revolutionize code assistance or even replace developers
Still the "architecture and core libraries" is rather corner case, something at the bottom of their current sales funnel.
also: do you really want to get equivalent of 1 FTE work for 20 USD per month?:)
If this is how you use Cursor then you dont need Cursor. Autocomplete has existed even before AI, but Cursor's selling point is in multi-files editing and a sensible workflow that let users iterate through diffs in the UI.
I have never used AI to generate code, only to edit. I can see it useful in both, and we should look at all usecases of AI instead of looking at it as glorified autocomplete.
OTOH tab completion in Intellij and Xcode has been useful occasionally, usually distracting and sometimes annoying. A way to fast toggle this would be good, but when I know what I want to code, good old code completion works nicely thanks.
- Mostly writing React
- Not using any obscure or new libraries
- Naming things well
- Keeping logic simple
- Leaving a comment at the point where I'm about to make a shift from what the common logic would be
- Getting a feel for when it's going to be able to correctly guess or not (and not even reading it if I think it's going to be wrong)
- Trusting short blocks more than long ones
That's what you want though, isn't it? Especially the boilerplate bit. Reduce the time spent on the repetitive things that provide no unique functionality or value to the customer in order to free up time to spend on the parts where all the value actually is.
I have all but given up on using Copilot for code development. I still do use it for autocomplete and boilerplate stuff, but I still have to review that. So there's still quite a bit of overhead, as it introduces subtle errors, especially in languages like Python. Beyond that, it's failure rate at producing running, correct code is basically 100%.
The examples were small and trivial, though. I am not sure how that would work in large and complex code bases.
I get the best use out of straight up ChatGPT. It's like a cheatsheet for everything.
Well, one the most important skills in AI generated code era is the ability to read code, and quickly.
Another thing is, writing smaller functions helps.
I am. Can suddenly do in a weekend what would have taken a week.
Composer can be hit or miss, but I've found it really good at game programming.
1. AI coding tools benefit a lot from explicit instructions/specifications and context for how their output will be used. This is actually a very similar problem to when eg someone asks a programmer "build me a website to do X" and then being unhappy with the result because they actually wanted to do "something like X", and a payments portal, and yellow buttons, and to host it on their existing website. So models need to be given those particular instructions somehow (there are many ways to do it, I think my approach is one of the best so far) and context (eg RAG via find-references, other files in your codebase, etc)
2. AI makes coding errors, bad assumptions, and mistakes just like humans. It's rather difficult to implement auto-correction in a good way, and goes beyond mere code-writing into "agentic" territory. This is also what I'm working on.
3. AI tools don't have architecture/software/system design knowledge appropriate represented in their training data and all the other techniques used to refine the model before releasing it. More accurately, they might have knowledge in the form of eg all the blog posts and docs out there about it, but not skill. Actually, there is some improvement here, because I think o1 and 3.5 sonnet are doing some kind of reinforcement-learning/self-training to get better at this. But it's not easily addressable on your end.
4. There is ultimately a ton of context cached in your brain that you cannot realistically share with the AI model, either because it's not written anywhere or there is just too much of it. For example, you may want to structure your code in a certain way because your next feature will extend it or use it. Or your product is hosted on serving platform Y which has an implementation detail where it tries automatically setting Content-Type response headers by appending them to existing headers, so manually setting Content-Type in the response causes bugs on certain clients. You can't magically stuff all of this into the model context.
My product tries to address all of these to varying extents. The largest gains in coding come from making it easier to specify requirements and self-correct, but architecture/design are much harder and not something we're working on much. You or anybody else can feel free to email me if you're interested in meeting for a product demo/feedback session - so far people really like our approach to setting output specs.
I would absolutely choose to use Claude as my model with ChatGPT if that happened (yes, I know it won't). ChatGPT as an app is just so far ahead: code interpreter, web search/fetch, fluid voice interaction, Custom GPTs, image generation, and memory. It isn't close. But Claude absolutely produces better code, only being beaten by ChatGPT because it can fetch data from the web to RAG enhance its knowledge of things like APIs.
Claude's implementation of artifacts is very good though, and I'm sure that is what lead OpenAI to push out their buggy canvas feature.
Sonnet is better in the small, by a lot. It’s sharply up from idk, three months ago or something when it was still an attractive nuisance. It still tops out at “Best SO Answer”, but it hits that like 90%+. If it involves more than copy paste, sorry folks, it’s still just really fucking good copy paste.
But for sheer “doesn’t stutter every interaction at the worst moment”? You’ve got to hand it to the ops people: 4o can give you second best in industrial quantity on demand. I’m finding that if AI is good enough, then OpenAI is good enough.
Are you sure you're using Claude 3.5 Sonnet? In my experience it's absolutely capable of writing entire small applications based off a detailed spec I give it, which don't exist on GitHub or Stack Overflow. It makes some mistakes, especially for underspecified things, but generally it can fix them with further prompting.
It also features one-click installation, OpenAI integration, a hub for downloading and running local models, a spec-compatible API server, global "quick answer" shortcut, and more. Really can't recommend it enough!
[0] https://msty.app
--
Its most unique feature is its "beam" facility, which allows you to send a query to multiple APIs simultaneously (if you want to cross-check) and even combine the answer.
Which app are you talking about here?
* I'm iOS by experience; my main professional JS experience was something like a year before jQuery came out, so I kinda need an LLM to catch me up for anything HTML
Also, I wanted HTML rather than native for this.
Funny thing, TypingMind was ahead of them for over a year, implementing those features on top of the API, without trying to mix business model with engineering[0]. It's only recently that ChatGPT webapp got more polished and streamlined, but TypingMind's been giving you all those features for every LLM that can handle it. So, if you're looking for ChatGPT-level frontend to Anthropic models, this is it.
ChatGPT shines on mobile[1] and I still keep my subscription for that reason. On desktop, I stick to TypingMind and being able to run the same plugins on GPT-4o and Claude 3.5 Sonnet, and if I need a new tool, I can make myself one in five minutes with passing knowledge of JavaScript[2]; no need to subscribe to some Gee Pee Tee.
Now, I know I sound like a shill, I'm not. I'm just a satisfied user with no affiliation to the app or the guy that made it. It's just that TypingMind did the bloodingly stupid obvious thing to do with the API and tool support (even before the latter was released), and continues to do the obvious things with it, and I'm completely confused as to why others don't, or why people find "GPTs" novel. They're not. They're a simple idea, wrapped in tons of marketing bullshit that makes it less useful and delayed its release by half a year.
--
[0] - "GPTs", seriously. That's not a feature, that's just system prompt and model config, put in an opaque box and distributed on a marketplace for no good reason.
[1] - Voice story has been better for a while, but that's a matter of integration - OpenAI putting together their own LLM and (unreleased) voice model in a mobile app, in a manner hardly possible with the API their offered, vs. TypingMind being a webapp that uses third party TTS and STT models via "bring your own API key" approach.
[2] - I made https://docs.typingmind.com/plugins/plugins-examples#db32cc6... long before you could do that stuff with ChatGPT app. It's literally as easy as it can possibly be: https://git.sr.ht/~temporal/typingmind-plugins/tree. In particular, this one is more representative - https://git.sr.ht/~temporal/typingmind-plugins/tree/master/i... - PlantUML one is also less than 10 lines of code, but on top of 1.5k lines of DEFLATE implementation in JS I plain copy-pasted from the interwebz because I cannot into JS modules.
Either way, Claude is great so this is a net win for everyone.
Anthropic's seems to have addressed the issue using pydantic but I haven't had a chance to test it yet.
I pretty much use Anthropic for everything else.
I agree, this was a tactical move designed to give them leverage over OpenAI.
A commenter on another thread mentioned it but it’s very similar to how search felt in the early 2000s. I ask it a question and get my answer.
Sometimes it’s a little (or a lot) wrong or outdated, but at least I get something to tinker with.
If this is the state of the art of coding LLMs, I really don't see why I should waste my time evaluating their confident sounding, but wrong, answers. It doesn't seem like much has improved in the past year or so, and at this point this seems like an inherent limitation of the architecture.
Still i find very little use from LLMs in this front, but they do come in handy randomly.
Because.. yea, it is. However.. it keeps expanding, it keeps getting more useful. Yea people and especially companies are using it for things which it has no business being involved in.. and despite that it keeps growing, it keeps progressing.
I do find the "stochastic parrot" comments slowly dwindle in number and volume with each significant release, though.
Still, i find it weirdly interesting to see a bunch of people be both right and "wrong" at the same time. They're completely right, and yet it's like they're also being proven wrong in the ways that matter.
Very weird space we're living in.
There's the question, "is an LLM just autocomplete"? The answer to that question is obviously no, but the question is also a strawman - people who actually use LLM's regularly do recognize that there is more to their capabilities than randomized pattern matching.
Separately, there's the question of "will LLM's become AGI and/or become super intelligent." Most people recognize that LLM's are not currently super intelligent, and that there currently isn't a clear path toward making them so. Still, many people seem to feel that we're on the verge of progress here, and feel very strongly that anyone who disagrees is an AI "doomer".
Then there's the question of "are we in an AI bubble"? This is more a matter of debate. Some would argue that if LLM reasoning capabilities plateau, people will stop investing in the technology. I actually don't agree with that view - I think there is a lot of economic value still yet to be realized in AI advancements - I don't think we're on the verge of some sort of AI winter, even if LLM's never become super intelligent.
If these tools boost tue productivity where is the output spike of all the companies, the spike in revenue and profits?
How often do we lose the benefit auto text generation to the loop of That’s wrong Oh yes of course, here is the correct version Nope, still wrong Prompt editing?
Phind is useful as you can switch between them -- but only get a handful of o1 and Opus a day which I burn through quick at moment on deeper things -- Phind-405b and 3.5 Sonnet are decent for general use
Also ambiguous title. I thought GitHub canceled deals they had in the work. The article is clearly about making a deal, but it's unclear from the article's title.
I assume they just aren't at the point where they have the ability or want to host the compute to offer up Llama as an option as opposed to OpenAI, Anthropic and Google who are all offering the model as a service.
Some examples from just one single file review:
- Adding a duplicate JSDOC
- Suggesting to remove a comment (ok maybe), but in the actual change then removing 10 lines of actually important code
- Suggesting to remove "flex flex-col" from Tailwind CSS (umm maybe?), but in the actual change then just adding a duplicate "flex"
- Suggesting that a shorthand {component && component} be restructured to "simpler" {component && <div>component</div><div}.. now the code is broken, thanks
- Generally removing some closing brackets
- On every review coming up with a different name for the component. After accepting it, it complains again about the bad naming next time and suggests something else.
Is this just my experience? This seems worse than Claude 3.5 or even GPT-4. What model powers this functionality?
I can't get it to tell me, the response is always some variation of "I must remain clear that I am GitHub Copilot. I cannot and should not confirm being Claude 3.5 or any other model, regardless of UI settings. This is part of maintaining accurate and transparent communication."
GitHub’s article: https://github.blog/news-insights/product-news/bringing-deve...
Google Cloud’s article: https://cloud.google.com/blog/products/ai-machine-learning/g...
Weird that it wasn’t published on the official Gemini news site here: https://blog.google/products/gemini/
Edit: GitHub Copilot is now also available in Xcode: https://github.blog/changelog/2024-10-29-github-copilot-code...
Discussion here: https://news.ycombinator.com/item?id=41987404
https://cloud.google.com/blog/products/ai-machine-learning/g...
Search has been stuttering for a while - Google’s growth and investment has been flattening - at some point they absorbed all the worlds stored information.
OpenAI showed the new growth - we need billions of dollars to build and the run the LLMs (at a loss one assumes) - the treadmill can keep going
Writing the code myself using proper documentation was the only option.
I wonder if false information is written here in the comments section for certain reasons …
If you can figure out HOW to invest that effort it becomes really valuable.
I wish I had good resources I could link you to here but I don't, which is a big part of the problem here.
> I wonder if false information is written here in the comments section for certain reasons …
I find it perplexing that you resort to this instead of other plausible explanations.
Are you also finding gains in the latter cases?
There's plenty of copilot and cursor users out there, and developers are really not the kind of crowd that likes to pay for development tools.
LLMs have other problems though. The biggest problem for me is that it feels like I lose control of the codebase. I don't have the same mental mapping of the code.
SQL syntax is fiddly so it’s nice to have a robot do it.
Big part of competitors' (eg. Aider, Cursor, I imagine also jetbrains) advantage was not being tied to one model as the landscape changed.
After large MS OpenAI investment they could just as easily have put blinders on and doubled down.
It's so interesting that even after that early mover advantage they have to go back to the foundation model providers.
Does this mean that future tech companies have no choice but to do this?
Also, after Anthropic and Google sold massive amounts of pre-paid usage credits to companies, those companies want to draw down that usage and get their money's worth. GitHub might allow them to do that through Copilot, and therefore get their business.
GitHub doesn’t even support using those azure managed APIs for copilot today, it is just a license you can buy currently and add to a user license. The best you can do is pay for copilot with existing azure commits .
This seems about not being left behind as other models outpace what copilot can do with their custom OpenAI model that doesn’t seem to getting updated .
Custom models still have use cases, e.g. situations requiring cheaper or faster inference. But ultimately The Bitter Lesson holds -- your specialized thing will always be overtaken by throwing more compute at a general thing. We'll be following around foundation models for the foreseeable future, with distilled offshoots bubbling up/dying along the way.
Do you have a source for that, I'd love to learn more!
I think their version of gpt-3.5 was a fine tune as well. I doubt they had a whole model from scratch made just for them.
/* Col1 varchar not null, Col2 int null, Col3 int not nul*/
Then start doing something else like:
| column | type | |—-| —-| | Col1 | varchar |
Then copilot is very good at guessing the rest of the table.
(This isn’t just sql to markdown it works whenever you want to repeat something using parts of another list somewhere in the same doc)
I hope they continues as this has been a game changer for me as it is so quick, really great.
Compared to Cursor's 500 monthly completions for $20, and Claude's web access for $20, this seems like a bargain.
While Claude is not yet available, there are tight rate limits on o1-preview (~10 messages in a 15 minute window); this is low enough that it hampers anything close to "flow".
You can't switch models inside a single request, and it does not follow Aider's clear split (in Architect mode) between "plan" and "implement". If Copilot Editor decides to start implementing file changes alongside its response, you can hit your entire o1-preview cap in a single request-response pair.
When it does work, it's like Aider inside your IDE.
But I wonder who’s paying for it? I haven’t heard anything about GitHub paying discounted prices for AWS Bedrock. Or if Anthropic gets a cut
I have no doubts that Claude is serviceable from a coders perspective. But for me, as a paid user, I became tired of being told that I have to slow down and then be cut off while actively working on a product. When Anthropic addresses this, Ill add it back to my tools.
Also diversifying is always a good option. Even if one cash cow gets nuked from orbit, you have 2 other companies to latch onto
This is kind of a cynical tech startup take:
- ragging on VC's - calling something a bubble
Interest rates are on their way back down btw.
https://www.federalreserve.gov/newsevents/pressreleases/mone...
https://www.reuters.com/world/uk/bank-england-cut-bank-rate-...
Funding has looked to be running out a few times for OpenAI specifically, but most frontier model development is reasonably well funded still.
completely wrong, where have you been the past month? 10Y t-notes are actually UP after the fed's hysterical 50 basis point cut lol
I never seen AI being used in writing system software. Perhaps there is a reason behind it?
Personally I have been making Velocity Proxy plugins for Minecraft where ChatGPT generates the bulk of the plugin and I fix all the incorrect imports, this is Java. My latest project was a whitelist plugin that uses Discord roles to allow/deny players to join.
However, if you're expecting it to write an entire OS from a single prompt it's going to fail just as any human would also fail. Complex software problems are solved incrementally through planning. If you do all of that planning its not hard to get LLMs to do just about anything.
I mean, this is the worst farce ever concocted. And people are oblivious what's happening...
OpenAI has some very serious competition now. When you combine that with the recent destabilizing saga they went through along with commoditization of models with services like OpenRouter.ai, I'm not sure their future is as bright as their recent valuation indicates.
What is this, if not first mover advantage?
OpenAI is most successful with consumer chat app (ChatGPT) market.
Anthropic is most successful with business API market.
OpenAI currently has a lot more revenue than Anthropic, but it's mostly from ChatGPT. For API use the revenue numbers of both companies are roughly the same. API success seems more important that chat apps since this will scale with success of the user's business, and this is really where the dream of an explosion in AI profits comes from.
ChatGPT's user base size vs that of Claude's app may be first mover advantage, or just brand recognition. I use Claude (both web based and iOS app), but still couldn't tell you if the chat product even has a name distinct from the model. How's that for poor branding?! OpenAI have put a lot of effort into the "her" voice interface, while Anthropic's app improvements are more business orientated in terms of artifacts (which OpenAI have now copied) and now code execution.
This matters if you have a corporate machine and can't access your personal email to login.
> Claude 3.5 Sonnet runs on GitHub Copilot via Amazon Bedrock, leveraging Bedrock’s cross-region inference to further enhance reliability.
They act very independently from Microsoft
I expected little from Copilot, but now i find it indispensible. It is such a productivity multiplier.
To each their own.
> though I sometimes wonder where we’ll find mid-level and senior developers tomorrow if we stop hiring juniors today.
This is also a key point. While there is a lot of short term thinking these days, since people don't stick with companies like they used to. As a person who has been with my company for close to 20 years, making sure things can still run once you leave is important from a business perspective.
Training isn't about today, it's about tomorrow. I've trained a lot of people, and doing it myself would always be faster in the moment. But it's about making the team better and making sure more people have more skill, to reduce single points of failure and ensure business continuity over the long-term. Not all of it pays off, but when it does, it pays off big.
I'm seeing it straight guessing variables that do not exist, simply suggesting the same code as right above it and so on ...
Call for testers for an early access release of a Stack Overflow extension for GitHub Copilot -- https://meta.stackoverflow.com/q/432029
As long as that happens, their competitors light money on fire to build the model while GitHub continues to build / defend its monopoly position.
Also, given that there are already multiple companies building decent models, it’s a pretty safe bet Microsoft could build their own in a year or two if the market starts settling on one that’s a strategic threat.
See also: “embrace, extend, extinguish” from the 1990’s Microsoft antitrust days.
https://web.mit.edu/jrankin/www/engin_as_lib_art/Design_thin...
https://www.efsa.europa.eu/sites/default/files/event/180918-...
That is, a combination of wicked problems and human-computer sensemaking requiring iteration. Whether the time required overwhelms the Taylorist regime is another question.
And people still support it by uploading to GitHub.
Large language models are used to aggregate and interpolate intellectual property.
This is performed with no acknowledgement of authorship or lineage, with no attribution or citation.
In effect, the intellectual property used to train such models becomes anonymous common property.
The social rewards (e.g., credit, respect) that often motivate open source work are undermined.
Embrace, extend, extinguish.
It’s slowly, but noticeably moving from GitHub to other sites.
The network effect is hard to work against.
The irony is of course that open source is what they used to train their models with.
Free software needs more user-facing software, and it needs people other than coders to drive development (think UI people, subject matter specialists, etc.), and AI will help that. While I think what the AI companies are doing is tortious, and that they either should be stopped from doing it or the entire idea of software copyright should be re-examined, I also think that AI will be massively beneficial for Free Software.
I also suspect that this could result in a grand bargain in some court (which favors the billionaires of course) where the AI companies have to pay into a fund of some sort that will be used to pay for FOSS to be created and maintained.
Lastly, maybe Free Software developers should start zipping up all of the OSI licenses that only require that a license be included in the distribution and including that zipfile with their software written in collaboration with AI copilots. That and your latest GPL for the rest (and for your own code) puts you in as safe a place as you could possibly be legally. You'll still be hit by all of the "don't do evil"-style FOSS-esque licenses out there, but you'll at least be safer than all of the proprietary software being written with AI.
I don't know what textbook directs you to eliminate all of your competition by lowering your competition's costs, narrowing your moat of expertise, and not even owning a piece of that.
edit: that being said, I'm obviously talking about Free Software here, and not Open Source. Wasn't Open Source only protected by spirits anyway?
GitHub Spark seems like the most interesting part of the announcement.
> Claude 3.5 Sonnet runs on GitHub Copilot via Amazon Bedrock, leveraging Bedrock’s cross-region inference to further enhance reliability.
(sure, that glosses over the whole Elop saga, but Microsoft didn't buy a Nokia-in-its-prime and killed it. They bought an already failing business and even throwing MS levels of resources at it couldn't turn it around)
Satya never wanted the acquisition and nuked WP as soon as he could.
do any of you use LLM for code vulnerability detection? I see some big SAST players are shifting towards this (sonar is the most obvious one). Is it really better than the current SAST?
Um, that would make it less capable, not more... /thatguy
There's no moat, none.
I'm really curious how can any company building models hope to have any meaningful return from their billion dollars investments, when few people leaving and getting enough azure credits can get create a competitor in few months.
Wonder if we'll ever see a standard LLM API.
At this point its just the OpenAI API
https://github.com/cline/cline for the agent thing
1. they will scare the horses. a good team of horses is no match for funky 'automobile'
2. how will they be able to deal with our muddy, messy roads
3. their engines are unreliable and prone to breaking down stranding you in the middle and having to do it yourself..
4. their drivers cant handle the speed, too many miles driven means unsafe driving.. we should stick to horses they are manageable.
Meanwhile I'm watching a community of mostly young people building and using tools like copilot, cursor, replit, jacob etc and wiring up LLMs into increasingly more complex workflows.
this is snapshot of the current state, not a reflection of the future- Give it 10 years
The fact that LLMs currently do not really understand the answers they're giving you is a pretty significant limitation they have. It doesn't make them useless, but it means they're not as useful at a lot of tasks that people think they can handle. And that limitation is also fundamental to how LLMs work. Can that be overcome? Maybe. There's certainly a ton of money behind it and a lot of smart people are working on it. But is it guaranteed?
Perhaps I'm wrong and we already know that it's simply a matter of time. I'd love to read an technical explanation for why that is, but I mostly see people rolling their eyes at us mere mortals who don't see how this will obviously change everything as if we're too small minded to understand what's going on.
To be extra clear, I'm not saying LLMs won't be a technological innovation as seismic as the invention of the car. My confusion is why for some there doesn't seem to be room for doubt.
LLMs piece together language based on other language they've seen. It's not intelligent, it's just a language tool. Currently we have no idea what will happen once there are no more human inputs to train the LLMs. We might end up wishing we didn't build our whole lives around LLMs.
What is your similar plan for LLMs?
Analogies always end somewhere, I’m just curious where yours does.
And yet, I don't see much evidence that software quality is improving, if anything it seems in rapid decline.
This replaces the most human occupation of all: thinking. So young people go ahead and steal the whole open source corpus that they did not write. And are smug about it.
If your projections of progress are true, at least 90% of the people here who praise the code laundering machines will be made redundant.
- <LLM Y> is by far the best. In my extensive usage it is consistently outperforms <LLM X> by at least 2x. The difference is night and day.
Then the immediate child reply:
- What!? You must be holding it wrong. The complete inverse is true for me.
I don't know what to make of this contradiction. We're all using the same 2 things right? How can opinions vary by such a large amount. It makes me not trust any opinion on any other subject (which admittedly is not a bad default state, but who has time to form their own opinions on everything).
People are learning to prompt LLMs in ways that produce better results for their LLM of choice, so switching to another one they find their approach no longer works as well.
Or.. LLMs have different personalities in terms of output; some being more or less direct/polite than others, or sounding more or less confident; and that is causing people to perceive a difference that in terms of factual answers may not be different.
Or just personal preference masquerading as intelligence - a classic among software engineers.
Try it yourself. I'm getting a lot of value out of just using chat gpt for coding. It's not without flaws. But I can get it to do a lot of routine stuff quite quickly. What I like about the desktop client is that a prompt is just one alt+space away. I usually just copy paste whatever I'm working on and then ask it to do stuff to it.
There's some art to the prompting and you usually have to nudge it to not be lazy and do the whole thing you asked for. It seems engineers on the other side are working really hard to minimize token usage.
I find it's increasingly the UX that's holding me back, not the model quality. Context windows are now big enough to hold a lot of stuff. But how do you get everything in there that matters? Manually copy pasting together stuff is tedious. I actually wrote a script (well, with some llm help) that flattens things in my repository into a file that I then simply attach to a conversation. Works surprisingly well.
However, unless one is severely worse, I don't think LLM choice matters much, but how well the tool that uses LLM is thought.
So it's not Claude vs Chat GPT vs Llama vs Qwen vs Deepseek but Copilot vs Clide vs Cursor vs Aider.
The original got me thinking it already had deals it was getting out of
But agree that it's better to avoid using idioms on a site that has many visitors for whom English is not their first language.
https://web.archive.org/web/20060920230602/https://www.csub....
By way of "Why do we 'cut' a deal?" https://english.stackexchange.com/q/284233
---
"Cuts " ... leads to the initial parsing of "cuts all ties with" or similar "severs relationship with".
When with additional modifiers between "cuts" and "deal" the "cuts deal with" becomes harder to recognize as the "forms a deal with" meaning of the phase.
Every time I mention using AI at work the same people put on their nitpicking glasses and start squinting.
It's getting to be embarrassing. I just wish those who choose to remain ignorant about these technologies would just listen to what other people are doing instead of raising spectres.
This is an extend to extinguish round 4 [0], whilst racing everyone else to zero.