First look at Microsoft 365 Copilot
paulrobichaux.com
paulrobichaux.com
I asked it a few simple fact based questions, such as who our CEO is. It answered with our previous CEO based on an old press release. I asked it what year it is, and it couldn’t say. A few other specific fact based questions were either wrong or unanswered.
So I asked it higher level questions about what the corporation does. It was able to answer those.
Finally, I began asking questions about policies in the handbook. When I asked about harassment, it refused to provide any answers. It felt like I was filtered. Filtering will kill the utility of these products. Something like a legal practice is messy business that the filter won’t like.
Finetuning should be reserved for desired behavior of the LLM, not to add new knowledge. For that, RAG is better suited.
(Is this about effectiveness, training time, forgetting, or something else?)
RAG combines a language model like GPT with a real-time search component. This allows the model to pull in information from external sources during its response generation process. Now the ability to access and integrate the most recent information is gained, which the language model alone wouldn't have.
You can think of an LLM as a set of basis vectors for human knowledge. If I feed in a PR training manual that is not in its dataset, it nevertheless figures out “hey I can make a reasonable approximation of this by combining X, Y, and Z” where X, Y, and Z are things it learned form its training set. In other words it maps the input into a vector representation based on its training data.
But in linear algebra two mappings can represent the same vector, just using different basis, so long as the vector space for the two basis are equal (or at least one is a subspace of the other). That's essentially what's going on here. A LLM builds a vector space on top of all human knowledge. If its parameters and training set are large enough, then the basis is in fact sufficient for representing anything you might throw at it. It will represent it in terms of it training set, yes, but that representation is high fidelity enough to represent the document in its entirety.
Fine-tuning a model is essentially rebalancing the initial weights of the LLM to pay special attention to certain clusters in its vector space, represented by the fine-tuning data. It's as if I threw random 2D points at a machine learning algorithm and it learned the basis { (1, 0), (0, 1) } representing the x-axis and y-axis. As a consequence of how inference works, it may then end up preferring to generate points when asked which are nearer to one axis or the other.
But then I fine-tune it on points that are distributed along the diagonal. This is not representative of the original training data, but NOT "outside" the original data. These points are fully represented by a linear combination of the x- and y- basis vectors. Nevertheless, the fine-tuning trains the model to prefer points which have weights that are multiples of (1, 1) or (1, -1) when represented in the original model. In other words, points along the diagonals.
Pragmatically speaking, this is no different from doing a whole new training run on the diagonal points, except that it is much, much cheaper, and has the capacity to reuse whatever knowledge was learned in the first training run.
"...all human knowledge..."
I am assuming in practice it was not trained on "ALL human knowledge", particularly in the case of our private company knowledgebase.Executives will ask questions about the technologies they hear about in the news, and it's good to be able to quickly speak to the good and the bad.
All sorts of different people may genuinely need those results: Attorneys, doctors and nurses, law enforcement, private investigators. Individuals searching for advice. Doing background checks on other people. Heck, just high school students researching a topic.
I'm not sure what to do about the (legitimate) concerns of the AI providers. They don't want a shitstorm, because someone got their AI model to [whatever]. But [whatever] can genuinely be important.
When filters are removed, one of the first things people do to the bots is make them racists. Time and time again. It happens every time.
The companies that make and operate these systems - like most companies - are more interested in avoiding PR problems than avoiding harm to humans. This is just one somewhat humorous example of that general policy.
What about pen companies? If somebody buys a Bic pen and writes a racist or hate-speech filled letter to somebody, should Bic be required to design a pen that censors or restricts hate speech? What's the underlying principle behind expecting a company to re-engineer their product to prevent any misuse?
I think most people agree that basic safety rails are desirable, such as child-resistant lids on medicine bottles, or guards on razor blades, but who's opinion is most relevant when it comes to where to draw the line?
A real-world case might be Plex. How much responsiblity does Plex have to ensure that users aren't using it to self-host pirated material? or CSAM? With engineering effort they could pretty easily collect data on everything everybody is watching, including fingerprinting and stuff. Should they be expected or required to do that?
Microsoft and others don’t want to facilitate that. I think it’s a reasonable concern. From a public policy perspective having millions of people who can instantly produce reams of racist abuse to swamp online fora is a problem, even if you believe in free speech. It potentially changes the power dynamic in favour of a racist minority to dominate public discourse.
Social proof and social learning have powerful effects on behaviour. LLMs are great, but used at scale they will have downsides too.
I kinda think when (not if) this happens it's going to be quickly drowned in all sorts of other AI generated garbage. In fact, we will probably render internet news useless pretty sure without a proper trust / source identification system.
I really think this will be the ultimate use case of blockchain tech/crypto.
You nearly always just see dog whistles, whataboutism, “Just Asking Questions”, “I’m not a racist but…”, fake stats, “racism is free speech”, etc. in fact, there’s quite a lot of that in this very section.
The question really is not “why does the model prevent sensitive topics”. The question is “why does hacker news get suddenly flooded with people who are ‘concerned’ about letting models be racist every single time models are discussed?”
I would suggest to you that there’s already loads of bots here.
Queries powered by Microsoft servers will be filtered.
Similarly, businesses are free to create their downloadable models however they so choose, and you are free to not use that model, preferring another one instead.
Frankly, it says a lot about the commenters here that the very first thing they test and judge a model on is whether or not they can force models to perform sensitive results.
So what? Seriously, so what?
Look, you can open Word on your system and type all sorts of racist stuff. Should Word auto-detect that (maybe with the help of AI) and refuse to accept your input? I don't think many people would find that to be a good idea.
Another point: Individual users don't affect the knowledge base. LLMs are not like Tay, dynamically integrating user input. An LLM like ChatGPT may adapt within a particular conversation, but that affects literally no one but you.
You jest but please don't give any Microsoft employees gleaning this thread any great ideas to show off in their next product meeting.
I didn't want to know that. I did not want to know that.
Censorship only exists because people are more lazy than they are nefarious. The other aspect is protecting idiots from themselves.
Whether that’s right or wrong, I’m not the judge.
Sure.
> If I have a blueprint for a homemade explosive that a 12 year old can make in a day, should that be easily accessible?
Yes.
Maybe as a society we then will have to have the conversations about cause/effect and personal responsibility that prepare children for adulthood.
When I was 9 my brother gave me a disc with all sorts of hacking and explosive making material. Jolly Rodger cookbook, handbooks on counterintel, hacking books about overflows.
I had great fun making thermite and smoke bombs and it led me to my career in software.
None of it was any more dangerous than climbing a tree because of good parenting.
Personal responsibility is great when it works with proper parenting. You’re going to have to deal with the real world when schools around America start blowing up to match the ideal.
Conservative religious nutjobs say the same shit when trying to ban books from libraries.
Everybody wants to ban the things they don't like, because most of us are weak-ass neophobes that bristle at the idea of thought that subverts these precious little needs. This is the whole problem.
We live in a diverse world. Stop coddling minds.
I remember reading books about the history of computer hacking and they started with stories of the phreakers and people who believed that all information should be free regardless of what humanity did with it. They always had stories about how they would pick the locks on the doors the labs that held the earliest computers. When staff found out and changed the locks, they simply scoffed and learned how to pick those locks instead. It wasn't until the school threatened them with expulsion did they finally relent.
In this day and age, as soon as you start putting up fencing around something, there's scores of people who just as anxious to find a way around it.
- What are my p/maternity entitlements?
- When will I be promoted?
- What will the bonus pool be this year?
- What will happen to me if my boss finds out I have been slacking off?
and many more examples of prompts that HR/legal simply will not accept a vague, or worse: incorrect, response to. Each of those types of questions require specific and considerate knowledge that LLMs simply cannot produce.
I wonder if it could be legally binding according to any actual law though.
While there are laws about minimum time in some states; the state is probably happy to enforce whatever extra time you grant. So if you tell your employees 20 they're probably going to be happy to hold you to 20.
What you're describing isn't a plausible situation. In the real world, conversations with real people have to happen when you take so much time off or have such a life changing event (especially with benefits implications). Obvious mistakes aren't generally legally binding, either. You can argue what isn't or isn't an obvious mistake, sure, but if you signed a contract that says 8 weeks, your boss tells you 8 weeks, there's posted documentation that says 8 weeks, etc., then saying you legitimately thought you had 20 weeks doesn't hold water.
If a company intends an HR bot to be relied on by employees, then its words should be legally binding. If the company doesn't intend for employees to rely on the bot, then it shouldn't exist, because nobody can make any plans based on its responses.
In most U.S. states your employment agreement is on an at-will basis and outside dismissal or discrimination for a set of protected reasons, you’re free to leave employment at any time, and your employer is free to let you go at any time.
Very few verbal interactions create a “binding commitment” for a company. All might be admissible if you brought suit, but “he said; she said; they said” is often more difficult with the passage of time (people are forgetful) and worse than written interactions, and just because you may be allowed to present them as evidence in your claim, that does not make anything a slam dunk legally.
A lot of what you will formally get upon receiving i.e. an offer of employment will be on company headed paper, signed by a person with authority (this is often “Delegated Authority”, that Role can do that Action is written down in Company Policies or a Company Handbook, which itself is reviewed and approved at very senior levels). That’s more “binding” and if you signed and returned, rejected other offers, then it fell through, you might see success in making a claim for losses and injury because you reasonably believed there would be a job and compensation. Even there it is going to say something to the effect of “this is not an employment contract, your status is at-will”, which significantly limits the liabilities of the company in most U.S. states.
So, short of the person you are talking to being a C-level executive or Company Officer, no, most of what you hear verbally isn’t “legally binding”.
Well, it was worse than useless as described in the premise by someone else as spitting out verifiably wrong information.
If a company intends an HR bot to be relied on by employees, then its words should be legally binding.
No company would would intend this, and it would be fun to see this "legally binding" status tested, because as my previous post and sibling comment lays out, it would be pretty hard to construct a plausible situation where a reasonable person believes it to be.
Like you pointed out, by preponderance of evidence, you could show that the employee should've reasonably understood they had "X" amount of days since its documented in so many different places and if one place said it was "XY" and employee decided to follow XY instead of X, they're fine since it was "legal binding". However, that would most likely not hold up in court because you can't use ignorance as a defense and say you the only vacation amount you knew of was what the BOT told you.
I think any employment attorney would inform their client this is what would happen and probably wouldn't take the case since the likelihood of winning would be so low.
Many states and many companies are “at will” employment.
As long as the termination isn’t for protected reasons you may not have much in the way of a civil law suit.
Ask me about the time I was “fired” for taking a vacation that my boss approved.
The day I returned, I was called into a meeting with HR and my boss, and berated for going on vacation without informing my boss (who again, approved my PTO), nor informing my team (a categorical lie). Later, my position was eliminated.
I think I need to start prompt injecting all my correspondence.
I’ve been trying to apply all the big name LLMs to analyze court opinions from the Caselaw project and holy mother of god, it’s all but useless in the most ridiculous way.
Criminal cases are almost universally guaranteed to trip up at least one of their safety filters. There’s some truly awful stuff in there… which is kinda the point of having an LLM read this stuff.
That and it hallucinates me a bunch of new case cites :D
this is the problem with basic RAG implementations that most products are currently using, simple vector search isn't able to handle queries like this.
The solution is enhancing the base query and also using structured data to filter on metadata
here's a good article covering some strategies - https://jxnl.github.io/instructor/blog/2023/09/17/rag-is-mor...
If the LLM is only fed the HR documents that have multiple dates, why would you expect it to answer that correctly? Why is it relevant that it answered it correctly.
I thought ".NET whatever" was confusing, and now it's Copilot.
In terms of the other Copilots, they all have different release dates.
Here's a topline review:
Copilot in Windows started rolling out on Windows 11 on September 26 through a Windows 11 update.
Copilot for Microsoft 365 began rolling out for enterprise customers on November 1 and will roll out to non-enterprise users at a later date.
Copilot for Sales will be available in the first quarter of 2024.
Copilot for Service will be generally available in early 2024.
Copilot in Viva will begin rolling out to customers "later in 2023", according to Microsoft.
Github Copilot was the first tool to release, all the way back in 2021.
https://www.zdnet.com/article/what-is-microsoft-copilot-here...
I dread to think how that is going to turn out for the rest of us.
My response to this is basically going to be to assuming that anything unsolicited is probably generated.
As it exists and is everywhere you can't just ban it away either. It is like steroids in baseball or adderall in med school, every other participate is cheating so you'll need to cheat just to keep up.
As you said it's like Adderall or steroids....I think it's more like the Internet, how many business in 2020 had zero Internet presence? What about 2000?
In the end we will get AGI and then super AGI which will be as full proof as a human could ever be with a team of researchers and fact checkers because it'll be genuine intelligence and able to know and discern more intuitively whatever it needs to know to get the job done.
Maybe AI never lives up to the hype, but a future conditional on AI transforming the world is not far from one destroying it.
AWS Copilot
Command line interface for containerized applications
So, I am not surprised that Copilot cannot answer information that is readily available in the right place. It is also unlikely that one can find that document, among many others, on the first try unless one is quoting a specific phrase.
The fundamental architecture needs to change. Copilot needs to act more like an agent - i.e., perform multi-step research to find this information and do it fast.
Doing it fast is another story, LLMs are pretty high latency at the moment.
Ex: Louie.AI will use a combination of tools to answer a question (database queries, RAG vector index lookups, Python, interactive charts, ...), and will often do multiple attempts on the same tool until it decides it has exhausted the immediate use of that tool.
Ex: Louie is even learning from usage, so as an analyst, if Louie stopped digging early and you decided to manually look further, Louie learns from this: It'll know in future sessions that it may be worth looking further in that kind of scenario.
None of that, in isolation, is unique to Louie.AI: It's just part of what it means to do a 'full' agent implementation vs a langchain/llamaindex/openai wrapper.
There's an interesting question around knowledge graph style questions here -- should an agent do iterative top-n vector similarity searches, or a single wider search over a knowledge graph, or maybe the documents should be combined into a knowledge graph and that's what's embedded. We're exploring a lot here in our bigger customer projects, and I can't say there's a clear universal answer...
From there on they use a special API that comes with Azure Open AI deployed GPT models that will look up using either cognitive search service or vector search.
That API is a black box, so either they use user message or they have the LLM write a search query.
I would assume 365 uses basically the same architecture.
[1] if you do the one click deployment and look through the deployed web app you will see it is based of this GitHub repo: https://github.com/microsoft/sample-app-aoai-chatGPT
The challenge Microsoft have set themselves however is that their problem domain is very large. They have to accommodate many different types of queries. We do benefit from having a specific type of usage pattern and much smaller data than a “real” Onedrive.
That said I am baffled that this product has performed as badly as it did in the review. These seem like pre-alpha type issues. I guess we should congratulate MS for being willing to put these things out to test at early stage but… come on!
I assume what you're meaning is that the generic model itself is not trained on company data, but is using retrieval techniques with the actual company data when used.
Retrieval Augmented Generation (apparently)
It's really helpful to define acronyms. HN has a broad readership - I follow AI and I had never come across this one.
RAG combines a language model like GPT with a real-time search component. This allows the model to pull in information from external sources during its response generation process. Now the ability to access and integrate the most recent information is gained, which the language model alone wouldn't have.
But even with RAG, results can disappoint. RAG works best on smaller chunks of information, but the knowledge in the underlying model gets in the way of accuracy. On our knowledgebase, in an area where there is a lot of mis- and not-quite-accurate info on the internet, it regularly provides inaccurate information—is unusable.
As @treprinium points out below, you also have to "calibrate similarity "thresholds" to know the probability distribution of relevant/correct/incorrect chunks for any given sentence and what the N should be to reach them. It's not going to be perfect but you can end up with 90% accuracy on average which tends to be better than many full-text search solutions"
I think if you survey AI at a high level you won’t encounter the term.
But if you survey or keep track with LLMs it’s an important concept that is widely known and discussed. It’s very important to know about because it’s one of the few practical techniques in LLMs.
Acronyms are only good for communicating phrases that you hear frequently. For infrequent terminology it's verging on puzzle-solving. Sometimes you can connect the dots and sometimes you can't.
I’ve found that learning an area solely by reading the literature etc is necessary but insufficient because you don’t get a sense of which topics are important. I’ve made this mistake several times in my career and ended up working on things that no one cared about.
Also, here’s a good primer on LLM concepts which should help in following discussions.
https://flyte.org/blog/getting-started-with-large-language-m...
It’s a bit meta but it’s also useful to read with the help of LLMs. For the past little while I’ve been reading in areas outside my expertise with the help of ChatGPT-4 defining and breaking things down for me as I read.
My comprehension speed has gone up tremendously and I can follow complex material without getting too lost.
And for me this is the dealkiller aganst BingGPT or however it is called these days. It just doesn't work if you compare it to a baseline of a simple query, which in addition has the benefit it will also directly give you the source and author of the statement. BingGPT will _try_ to disclaim the sources it uses but fail miserably, as a lot of the times the assertion will not be in the source, or the _opposite_ assertion will be in the source.
* Asking Power BI to analyze some employee data and generate a Power BI report page
* Asking Power BI to re-theme the page to match a different Power BI report
* Asking Power BI to refine some of the page content to answer a particular question about retention
* Asking Word to generate draft minutes from a Teams meeting transcript
* Asking Power Apps to create an app and iteratively build out features
* Asking Power Automate to map out an automation flow
This article seems to be focused on a very narrow understanding of Copilot's capabilities. I'm waiting to see if the above use cases are fact or fiction.
The not so fine print is Copilot may return garbage and it's up to you to confirm the results.
Its customer data, and it can technically be anything. Imagine making a system that is expected to be reliably excellent at generating things within that domain. It is already a very hard problem that I can confidently say is unsolved, and probably not even as good as a human can sort things. Even in a good dataset that is unstructured and unlabelled you will face issues in just a single domain (e.g. vision tasks trying to classify a variety of vehicles), imagine that being extended to a wide variety of subjects that can mix and match.
I imagine at most, it'll be a reputation hit.
Like, we get it - you resent your coworkers & consider them beneath you.
Hopefully, your unique, elite output keeps you employable for a few quarters longer than them.
I’m routinely asked to make reports, but things stall when I ask, “so what do you actually need to see on the report, and where should the data come from”. A lot of people out there are kind of cargo cutting their way through things.
Cargo culting. Though cargo cutting would be interesting to watch :-)
Your response was both snarky and personal which I think comes across much worse.
You would be amazed how many issued we identified so far and our mgmt uncharacteristically wants us to bring up issues. Some companies kill their messengers.
So yeah.. I absolutely buy that a portion of business reports are of low quality.
But if I was making mistakes, say, in 5% of my facts and figures...I probably would have been reassigned to other work. 5% is way too high for an Excel sheet: it makes the model completely useless. It often takes more effort to audit and fix 5% of busted Excel cells than it does to just rebuild the model from scratch. GPT-4/Copilot/Gemini/etc almost certainly make more than 5% errors for any real use case.
I am sure there are many bad organizations with lower standards than the organizations I worked at. My concern is that good organizations will adopt technology like GPT, inadvertently lowering their standards because they trusted a machine they didn't understand. It's very worrying that good people are rationalizing this choice by saying "well humans are also unreliable" without putting specific numbers to it. 5% errors in document summarization is not good enough for human workers, and it certainly shouldn't be good enough for software.
I will add that Robichaux is far more diligent and skeptical of Copilot than most people who use it, including Microsoft. MSFT's public Bing demo had a ton of awful mistakes, which is shocking: I assumed MSFT would have faked the demo! The fact that they didn't suggests they have a magical and undeserved belief in OpenAI's technology.
I suspect the demo you watched also had a ton of mistakes that went unnoticed - if Power BI makes a pretty graph with a plausible trendline, how many people are really going to go back and check the underlying code, making sure it's doing the right thing to the right data sources? If it generates draft minutes from a meeting transcript, but screws up 5% of the items, those errors likely won't be noticed in a quick spot-check, and will end up misleading people who didn't attend the meeting. Robichaux's post brings up a ton of issues that affect any use case - the tool just is not reliable.
In my experience with copilot for power automate, the demos you see are mostly smoke and mirrors. I tried to recreate the demo as exactly as I could, I couldn't. Maybe they can help absolute beginners, but if you know the first thing about power automate, you're better off with a web search than asking the copilot for help. This is still v1 though (if it's not still technically a "preview"), but it really should be marketed so hard because it is currently not delivering.
Out of those use cases, though, four of them are just using Copilot as a natural-language front end (e.g. re-theming in Power BI). I would expect that to work well and, frankly, don't think it's that interesting as a use case. Sort of like Copilot for Windows… why do I want a natural-language system to tell me where setting X lives?
It’s been rather good at the “office” side of things, and not very useful at doing the things we were doing with OpenAI tools. We too have found it rather good at summarising meetings, a very useful feature for any sort of Microsoft Teams meeting and much better than having some unlucky member of the meeting do it. Copilot has really shined for us, however, have been with PowerPoint, presentations in general and all sort of investor targeted material. I work in the energy sector, and we can’t exactly produce presentations and materials like a design heavy organisation might. So there are two consequences of this, one is that any internal presentation now looks good, which is basically a win-win for everyone. The other is how we’ve basically cut our external orders on design projects to a zero, which is a win did us and terrible for an entire business of graphic designers who target organisations like ours.
I’ve personally mostly used Copilot to make funny pictures after it became apparent that it wasn’t anywhere near the level of OpenAI products as far as behind useful in programming and technology goes. Well, unless it’s very Microsoft documentation related, then Copilot does well at pointing you in the right direction. Well, on the frontend side of things, it’s replaced all our payments to icon libraries, as it’s capable of producing all the icons we need very well. OpenAI can do this as well, but at a cost.
As far as the price goes. Copilot comes with our regular Azure and Office365 subscriptions, and similar to Microsoft teams it’ll be the obvious choice because of this. How will you ever justify paying for a competitor when your budgets show that you’re getting copilot for free? Well, you won’t in most organisations. I’m not a fan of this sort of monopoly, but it’s not like it will change unless EU or US regulation stops it.
One example was how our p1v3 ($0.17 listed) was much cheaper than our p1v2 ($0.11 listed) because our 3rd party agreement was set up that way. The most beautiful part about that, and these things in general, that nobody told our development teams this until around 3 months before we changed 3rd party vendor. A switch we also weren’t informed about, so we went from p1v2 to p1v3 and back again over a few months.
Similarly things like Microsoft teams, crazy amounts of SharePoint online document storage, PowerApps, 365 Copilot and so on, necessarily an increase in cost because they are included on the user licenses we pay per user which “bundles” things.
I can’t get much more technical on the actual pricing and licensing because it’s never been my field of expertise, and to be perfectly honest, I don’t really care about it. But things aren’t always as clear cut as the public price listings. When you drop a lot of money with Microsoft or Amazons AWS you get “features” that aren’t available to you or me.
It’s important to note that it’s not necessarily free just because it looks free on the budget, but budgets in enterprise organisations are a whole story of themselves.
https://www.npr.org/2023/12/14/1219246964/cobalt-is-importan...
In some sense, these companies are ultimately going to be the ones to make that decision just by providing these little helpful suggestions to billions of people.
Businesses want more features like this. If one employee writes something that hints that the employee reading it should do something, that might work in Japan, but in many other countries, a feature that suggests ways to make the action item clearer to the recipient would be well-received.
(Emphasis mine)
Yeah, its pretty easy to see why this would be desirable for some parties. But multiple things can be true at once.
It is a way to help out businesses, but its not just that. Its also a thing that is going to shape our language, which will in turn shape our thought and discussions. I just think we should be cognizant of who has their hands on those levers.
When you say 'just', do you mean you think this is the only effect?
As in the one and only effect of one of the most used text editors suggesting changes to its users writing will be to make their writing more clear? Thats quite optimistic.
If anything, it will make their writing more anodyne, which may be great from a business perspective. But it will probably have all sorts of effects that we won't be able to measure until years after the fact, if ever.
This is like Facebook making little tweaks to their timeline. Just by virtue of the fact that billions of people are exposed to it, you are actually shaping thought around the world, intentionally or not, knowingly or not.
This has been my experience with a lot of my interactions using ChatGPT.
The popular narrative about genAI is this ideal world where you just ask It a question and it gets you the “correct” answer every single time. The more I use it, unfortunately that is turning out NOT to be the case. I am still very optimistic about AI but this generation of AI feels like it has ways to go before it can be trusted
I think GitHub Copilot made this very obvious. It can outline stuff really well, but you really have to double check and fix everything that it spits out as soon as complexity is above a certain (very low) threshold.
ChatGPT is quite far from being the Google replacement that some people pretends/wants it to be. It's like that annoying friend who is extremely confident in announcing facts, but is often wrong. I keep giving it the side eyes.
I wanted to test Copilot using some actual data, but not data from work. The docs I loaded are a mix of PDFs and Office documents, with embedded graphics, tables, and so on. Asking questions about e.g. stall speed is a proxy for asking fact-based questions from a corpus of "real" work docs.
Some of the features I'm most excited about trying, like meeting recap, aren't available to me unless I schedule a bunch of meetings in the sandbox, which I can't do because it's a sandbox where I don't have access to invite outsiders.
I will say that the integration and fit/finish of Copilot throughout M365 is quite good. For the most part, it's very easy to discover the entry points and get Copilot to do stuff. When its results are good, it's a very useful tool… but MS has some overall work to do to get consistency on that point.
I don’t have access yet to the AI extensions in Microsoft 365 but I have spent a fair amount of time with Bard + Workplace plugins. The best experience has been giving Bard a link to recipes on YouTube and getting short text summaries that are adequate for making the recipe. Sometimes queries against my GMail and Docs are useful but still a work in progress.
Anyway, I have very high expectations that within 1 year, both Microsoft and Google will have nailed CoPilot-like AI tools for their cloud productivity apps.
Most people I know who work in large enterprises work in a "deck culture" where everything needs slides, and - again anecdotally - no one likes making slides.
For many people it's a slow and painful process. If Copilot 365 can take a Word doc that someone wrote and mash it with a corporate slide template to produce a usable, or near-usable, deck that just needs a few tweaks before being ready to circulate or present, that will be a huge win for so many people.
I saw a video demo of this recently and it was really impressive. A raw meeting transcript was turned into a meeting summary and that summary was then turned into nice-looking slides based on an existing presentation. All in under 5 minutes. I'm sure it was a controlled environment (the demo was by a Microsoft evangelist), but I didn't actually expect it to be that good.
As more and more of these AI assistants appear, I think they will all slowly find their niche like this.
I'm not sure about that. I think there's a lot of people who actually like frequent long meetings and bad slides. That's why we have them everywhere. The rest is groupthink.
I agree your use cases are the big wins though.
It is constantly misunderstanding my emails and email threads.
The generative part of it is worse in a way, I use it as a joke with friends, the output is corporate speak^10.
I do find it useful when emailing ppl who appreciate obfuscated long emails, I can get company support ppl to treat me well...
Interestingly ChatGPT-4 gets it right. I tried your first question with ChatGPT-4:
User: what’s the single-engine service ceiling of a Baron 55
ChatGPT: The single-engine service ceiling of a Beechcraft Baron 55 is approximately 7,000 feet. This is the maximum altitude at which the aircraft can maintain a specified rate of climb, usually 100 feet per minute, with one engine inoperative.
I have yet to see that not work using 3.5, though it doesn’t mean it won’t hallucinate if the injected text does not have the answer to the question.
As a side-note, OpenAI does very well on this currently in my experience. This morning, it even used search for a Rust deprecation warning that I got on some code and gave me a link to a related issue on a specific package that I was using. Quite amazing if you think about it since it (1) correctly understood the warning, (2) correctly decided to use search (Bing), (3) correctly wrote a search query, and (4) correctly gave the link to the right issue.
The total inability of transformer LLMs to summarize documents is not a surprise. Document summarization is an incredibly difficult task because it's impossible to do it accurately if you don't understand the semantics of the document. Current LLMs are simply too stupid to have anything but shallow, illusory understanding of human language. The illusion works when you are "kicking the tires" with simple documents, because literally thousands of data contractors have diligently trained the LLM on simple documents. But the illusion badly fails when you feed the LLM interesting documents:
- novel ideas are smashed against the training set and flattened into a homogenous mush
- specific facts and numbers are replaced with """statistically likely""" facts and numbers, which actually works okay for giant tables of statistics (except it destroys meaningful outliers). But this badly fails for, say, engineering specifications around aircraft engines. Or quarterly profit/loss figures, which Bing's demo screwed up.
- Even code generation shits the bed if you're working in an uncommon language. Last I checked, GPT-3.5 had a terrible understanding of F# and would copy-paste dozens of lines verbatim from specific Github projects, even if that code was inappropriate for the task at hand.
This illusion of competence is not meaningless if it works: GPT's Python code generation mostly works by translating English sentences into Python sentences, supplemented by plagiarism, but it's a useful tool regardless. The problem is that Copilot's illusion of competence does not work for aircraft piloting. There are not millions of lines of written aircraft logic in Github like there is with Python.
The only way I can see this tool working is if Microsoft goes back and pretrains Copilot specifically on aircraft specifications, then hires data contractors for RHLF specific to Robichaux's use case. Otherwise it's simply too unreliable.
I am less interested in it writing stuff for me as I can do that with chat-gpt.
A basic search engine search, while much more labor intensive, at least quickly reveals when there’s some inconsistency around a potential topic that allows an inquisitive individual to probe further. These new tools could be super dangerous in that they just spread false information and many folks don’t know or care enough to check if the answers are even right.
We’re going to see a lot more train-wrecks like that lawyer that used ChatGPT to do research on his brief and ended up citing a bunch of cases that were made up and didn’t exist.