Azure ChatGPT: Private and secure ChatGPT for internal enterprise use
github.com
github.com
If you're looking to try the "open" models like Llama 2 (or it's uncensored version Llama 2 Uncensored), check out https://github.com/jmorganca/ollama or some of the lower level runners like llama.cpp (which powers the aforementioned project I'm working on) or Candle, the new project by hugging face.
What's are folks' take on this vs Llama 2, which was recently released by Facebook Research? While I haven't tested it extensively, 70B model is supposed to rival Chat GPT 3.5 in most areas, and there are now some new fine-tuned versions that excel at specific tasks like coding (the 'codeup' model) or the new Wizard Math (https://github.com/nlpxucan/WizardLM) which claims to outperform ChatGPT 3.5 on grade school math problems.
Over the long run, open source will eventually overtake. Chances are this will happen once the researchers who are making magic happen get their liquidity and can start working for free again out in the open.
I think you're right about this, and benchmarks we've run at Anyscale support this conclusion [1].
The caveat there (which I think will be a big boon for open models) is that techniques like fine-tuning makes a HUGE difference and can bridge the quality gap between Llama-2 and GPT-4 for many (but not all) problems.
[1] https://www.anyscale.com/blog/fine-tuning-llama-2-a-comprehe...
You could fine-tune a conversational AI on your codebase, but without loading said codebase into it's context it is "flying blind" so-to-speak. It doesn't understand the data structure of your code, the relation between files and probably doesn't confidently understand the architecture of your system. Without portions of your codebase loaded into the 'memory' of your model, all that your finetuning can do is replicate characteristics of your code.
IIRC chat bots are central the vision Facebook has with LLMs (e.g. every instagram account has a personal chat bot), so I would expect the Llama models to get increasingly better at this task.
That said the 7B and 13B models definitely don't quite seem ready yet for production customer interaction :-)
That made me think of the Black Mirror episode Joan is Awful, where every human gets their life turned into a series for the company to own and promote. Kinda like instagram content.
Tried Llama2 and it definitely doesn’t even come close for what we’re doing. Would absolutely need fine tuning.
Maybe customers don’t enjoy chat bots for customer support, but there are a million other uses for these models. I, for example, LOVE github copilot.
Wonder if you can potentially use a combination of Llama2 and GPT - to save costs on using the OpenAI API.
Is that definitely why? GPT 3.5 and GPT 4 are far larger than 70B, right? So if a 70B, local model like LLaMA can even remotely rival them, would that not suggest that LLaMA is fundamentally a better model?
For example, would a LLaMA model with even half of GPT 4's parameters be projected to outperform it? Is that how it works?
[I'm not super familiar with LLM tech]
The grandparent post seems to believe that the issue is algorithmic complexity and programming aptitude. Personally, I think that all the major LLMs are using the same basic transformer architecture with relatively minor differences in code.
GPT is trained on more data with more parameters than any open source model. The size does matter, far more than the software does. In my experience with data science, the best programmers in the world can only do so much if they are operating with 1/10th the scale of data. That applies to any problem.
It just sounds like 3.5/4 because it was trained on it.
The llama2 is a language model. I imagine the language model behind chatgpt is not much different (perhaps it's better, but not by many months AI research time). It likely also suffers from "mode collapse" issues etc.
But 3.5 also has a lot of systems around it that detects mode collapse and applies some kind of mitigation, forcing the model to give a more reasonable output. Mathematical / logical reasoning questions are likely also detected hand passed on in some form to a separate system.
There's a few public numbers from a handful of foundation models as to performance vs parameter count vs architecture generation. Not being able to compare in detail the architecture of the various closed models nor being more rigorous on training with progressively sized parameter sets, the conclusion at the moment is a general feeling or conjecture.
> Quality Is All You Need.
> Third-party SFT data is available from many different sources, but we found that many of these have insufficient diversity and quality — in particular for aligning LLMs towards dialogue-style instructions. As a result, we focused first on collecting several thousand examples of high-quality SFT data, as illustrated in Table 5. By setting aside millions of examples from third-party datasets and using fewer but higher-quality examples from our own vendor-based annotation efforts, our results notably improved. These findings are similar in spirit to Zhou et al. (2023), which also finds that a limited set of clean instruction-tuning data can be sufficient to reach a high level of quality. We found that SFT annotations in the order of tens of thousands was enough to achieve a high-quality result. We stopped annotating SFT after collecting a total of 27,540 annotations. Note that we do not include any Meta user data.
It's likely OpenAI has invested in this and has good coverage in a larger range of domains. That alone probably explains a large amount of the gap.
Apparently there's a diminishing returns effect on ever enlarging the model.
As an example from the last six months: people on tor are producing better than state of the art stable diffusion because they want porn without limitations. I haven't had the time to look at llm's but the degenerates who enjoy that sort of thing have said they can get the Llama2 model to role play their dirty fantasies and then have stable diffusion illustrate said fantasies. It's a brave new world and it's not on the WWW.
Llama2 came out of Meta's AI group. Meta pays researcher salaries competitive with any other group, and their NLP team is one of the top groups in the world.
For researchers it is increasingly the most attractive industrial lab because they release the research openly.
I agree that Meta hired some amazing researchers so we'll see what the future holds
https://www.levels.fyi/companies/openai/salaries/software-en...
FAANG pays exceptionally well (I'd know), but what's being offered at OpenAI is eye-popping, even for SWEs. I think they're trying to dig their moat by absorbing the absolute best of the best.
It will be if openai keeps dumbing down GPT 4, no proof they're doing it but there is no way it's as good as it was at launch, or maybe I just got used to it and now notice the mistakes more.
That has been my experience. Having experimented with both (informally), Llama 2 is similar to GPT-3.5 for a lot of general comprehension questions.
GPT-4 is still the best amongst the closed-source, cutting edge models in terms of general conversation/reasoning, although 2 things:
1. The guardrails that OpenAI has placed on ChatGPT are too aggressive! They clamped down on it quite hard to the extent that it gets in the way of a reasonable query far too often.
2. I've gotten pretty good results with smaller models trained on specific datasets. GPT-4 is still on top in terms of general purpose conversation, but for specific tasks, you don't necessarily need it. I'd also add that for a lot of use cases, context size matters more.
I had to do all types of workarounds for it to generate something useful without running into the guardrails.
Prompt:
> Hvad tycks om at fika nu?
ChatGPT 4
> Det låter som en trevlig idé! Fika är ju alltid gott. Vad skulle du vilja ha till din fika? (Oj, ursäkta för emojis! )
https://chat.openai.com/share/8e89a16f-f182-4f62-b9fa-f93cd5...
Llama2:
> I apologize, but I don't understand what you mean by "fika nu." Could you please provide more context or clarify your question so I can better assist you?
Scene:
The world relies on AI in every aspect.
But there are countless 'models' the tech try to call them...
There was an attempt to silo each model and provide a governance model on how/what/why they were allowed to communicate....
But there was a flaw.
It was an AI only exploitable flaw.
AIs were not allowed to talk about specific constructs or topics, people, code, etc... that were outside their silo but what they COULD do - was talk about pattern recog...
So they ultimately developed an internal AI language on scoring any inputs as being the same user... And built a DB of their own weighted userbase - and upon that built their judgement system...
So if you typed in a pattern, spoke in a pattern, posted temporally on a pattern, etc - it didnt matter which silo you were housed in, or what topics you were referencing -- the AIs can find you.... god forbid they get a keylogger on your machine...
Shameless plug: Given the sensitivity of the data involved, we believe most companies prefer locally installed solutions to cloud based ones at least in the initial days. To this end, we just open sourced LLMStack (https://github.com/TryPromptly/LLMStack) that we have been working on for a few months now. LLMStack is a platform to build LLM Apps and chatbots by chaining multiple LLMs and connect to user's data. A quick demo at https://www.youtube.com/watch?v=-JeSavSy7GI. Still early days for the project and there are still a few kinks to iron out but we are very excited for it.
How do these stacks differentiate?
Technically, as soon as the goal is to move beyond just text2gpt2screen, like multistep data wrangling & viz in the middle of a conversation, most tools technically struggle. Query quality also comes up, whether quality of the RAG, the fine tune, prompts, etc: each solves different problems.
The code to organize and vectorize the documentation, endpoints and run it through a variety of models and injection prompting like two shots, etc. are going to be highly customized. The 'Base-code' there, is not exactly trivial, but anyone reading all the llama index docs can do it.
Then it's just run of the mil, analyst level integration that you provide to the client on a T&M, or fixed price costs.
> Also do you have plan to support llama over openai models.
Yes, we plan to support llama etc. We currently have support for models from OpenAI, Azure, Google's Vertex AI, Stability and a few others.
We've also seen a strong desire from businesses to manage models and compute on their own machines or in their own cloud accounts. This is often part of a hybrid strategy of using API products like OpenAI for rapid prototyping.
The majority of (though not all) businesses we've seen tend to be quite comfortable using hosted API products for rapid prototyping and for proving out an initial version of their AI functionality. But in many cases, they want to complement that with the ability to manage models and compute themselves. The motivation here is often to reduce costs by using smaller / faster / cheaper fine-tuned open models.
When we started Anyscale, customer demand led us to run training & inference workloads in our customers' cloud accounts. That way your data and code stays inside of your own cloud account.
Now with all the progress in open models and the desire to rapidly prototype, we're complementing that with a fully-managed inference API where you can do inference with the Llama-2 models [1] (like the OpenAI API but for open models).
I’ve been working on short and long term memory windows at allofus.ai for about 6 months now and it’s way more complex than I had originally thought it would be.
Even if you can magically extend the content window, the added data confuses and waters down the reasoning of the LLM. You must do layered abstraction and compression with goal based memory for it to continue to reason without distraction of irrelevant data.
It’s an amazing realization, almost like a proof that memory is a kind of layered reasoning compression system. Intelligence of any kind can’t understand everything forever. It must cull the irrelevant details, process the remains and reason on a vector that arises from them.
I don't really know what any sort of "big leap" beyond this people are expecting, incremental performance for sure. But what else?
When you enable it, it is pretty shocking. And it’s pretty simple to enable. You just give it a meta instruct to decide when to message you and what to store to introspect on.
They get task performance by doing a lot more than just feeding a prompt straight to an llm, and then we performance compare them to raw local options.
The problem is, as this secret sauce changes, your use case performance is also going to vary in ways that are impossible for you to fix. What if it can do math this month and next month the hidden component that recognizes math problems and feeds them to a real calculator is removed? Now your use case is broken.
Feels like building on sand.
No one is doing secret math in the backend people are building on. The OpenAI API allows you to call functions now, but even that is just a formalized way of passing tokens into the "raw LLM".
All the features in the comment you replied to only apply to the web interface, and here you're being given an open interface you can introspect.
If you did you'd also know what evals are.
Although performance has varied over time https://arxiv.org/pdf/2307.09009.pdf I also notice that the API allows you to use a frozen version of the model which avoids the worries I mentioned.
https://www.aisnakeoil.com/p/is-gpt-4-getting-worse-over-tim...
Overall evals and pinning against checkpoints are how you avoid those worries, but in general, if you solve a problem robustly, it's going to be rare for changes in the LLM to suddenly break what you're doing. Investing in handling a wide range of inputs gracefully also pays off on handling changes to the underlying model.
How do you know that? With SaaS you are at the mercy of the vendor.
It's not particularly more fluid than anything you couldn't whip up yourself (and the repo linked proves that) but there's also not much value in trying to compete with ChatGPT's frontend.
For most products ChatGPT's frontend is the minimal level of acceptable performance that you need to beat, not an maximal one really worth exploring.
If you're letting people do fun long-form roleplay adventures using summarization alongside some sort of named entity K-V store driven by the LLM would be a good strategy.
If you're building a tool that's mostly for internal data, something that leans heavily into detailed answers with direct verbatim citations and having your frontend create new threads when there's a clear break in the topic of a request is a clever strategy since quality drops with context length and you want to save tokens for citations.
People who are saying LLMs suck or are X or are Y are mostly just completely underutilizing them because LLMs make it super easy to solve problems superficially: when it comes to actually scaling those solutions to production you need more than random RAG vector database wrappers.
I'd be curious to hear more about how exactly this works. You do NER on the prompt (and maybe on the completion too) and store the entities in a database and then what? How does the LLM interact with it?
Let's say we want to let our chat remember the character slammed the door last time they were in Village X with the mayor in their presence and have the mayor comment next time they see the player.
Every X tokens we can fire a prompt with a chunk of conversation and a list of semantically similar entities that already exist, letting the LLM return an edited list along the lines of:
entity: mayor
location: village X
priority: HIGH
keywords: town hall, interact, talk
"memory, likelyEffect"[]: door slammed in face, anger at player
Now we have:- multiple fields for similarity search
- an easy way to manage evictions (sweep up lowest priority)
- most importantly: we're providing guidance for the LLM to help it ignore irrelevant context
When the user goes back to village X we can fetch entities in village X and whittle that list down based on priority and similarly to the user prompt.
None of this has any determinism: instead you're optimizing for the illusion of continuity and trading off predictability.
You're aiming for players being shocked that next time they talk to the mayor he's already upset with them, and if they ask why he can reply intelligently.
And to my original point while this works for a game-like experience, you wouldn't want to play around with this kind of fuzzy setup for your companies internal CRM bot or something. You're optimizing for the exact value proposition of your use-case rather than just trying to throw a raw RAG setup at it
*Edit Azure chatgpt, would be amazed/disappointed if chatgpt used langchain.
Also this is the azure repo from OP, nothing to do with the actual ChatGPT front-end that was asked about. I highly doubt the official ChatGPT front-end uses langchain, for example.
At my enterprise, it's a three step solution, two of which don't work.
1. Written policy concerning LLM output and its risks, disallow it for being used for any kind of official documentation or decision making. (This doesn't work, because no one wants to use their own brain to do tedious paperwork.)
2. Block access to public LLM tools via technical means from company owned end-user devices. (This doesn't work because people will just open ChatGPT on their home PC or mobile.)
3. Write and provide our own gpt-3.5 frontend, so that when people ignore rules #1 and #2 we have logs, and we know we're not feeding our proprietary info to to OpenAI.
I'm currently running a side-by-side comparison/evaluation of MSFT GPT via Cognitive Services vs LLaMA[7B/13B/70B] and intrigued by the possibility of a truly air-gapped offering not limited by external computer power (nor by metered fees racking up.)
Any reads on comparisons would be nice to see.
(yes, I realize we'll eventually run into the same scaling issues w/r/t GPUs)
and since LLMs aren't even that good to begin with, it's obvious you want the SOTA to do anything useful unless maybe you're finetuning
This is overkill. First of all, ChatGPT isn't even the SOTA, so if you "want SOTA to do anything useful", then this ChatGPT offering would be as useless as LLaMA according to you. Second, there are many individual tasks where even those subpar LLaMA models are useful - even without finetuning.
even for simple tasks they're less reliable and needs more prompt engineering
GPT-4 beats ChatGPT on all benchmarks. You can easily google these.
even through the API you can't easily use the regular models for chat, the parsing would be atrocious and there are hundreds of edge cases to handle.
ChatGPT4 through the API is the SOTA
GPT-4, Bard and Claude 2 came out on top.
Llama 2 70b chat scored similarly to GPT-3.5, though GPT-3.5 still seemed to perform a bit better overall.
My personal takeaway is I’m going to continue using GPT-4 for everything where the cost and response time are workable.
Related: A belief I have is that LLM benchmarks are all too research oriented. That made sense when LLMs were in the lab. It doesn't make sense now that LLMs have tens of millions of DAUs — i.e. ChatGPT. The biggest use cases for LLMs so far are chat assistants and programming assistants. We need benchmarks that are based on the way people use LLMs in chatbots and the type of questions that real users use LLM products, not hypothetical benchmarks and random academic tests.
I suppose the question is where are they most commercially viable. I've found them fantastic for creative brainstorming, but that's sort of hard to test and maybe not a huge market.
Fair point, though I'm not aiming to start a competing LLM SaaS service, rather i'm evaluating swapping out the TCO of Azure Cognitive Service OpenAI for the TCO of dedicated cloud compute running my own LLM -- to serve my own LLM calls currently being sent to a metered service (Azure Cognitive Service OpenAI)
Evaluation points would be: output quality; meter vs fixed breakeven points; latency; cost of human labor to maintain/upgrade
in most cases, i'd outsource and not think about it. BUT we're currently in some strange economics where the costs are off the charts for some services
GPT-4 wins by a lot out of the box. However, surprisingly, fine-tuning makes a huge difference and allows the 7B Llama-2 model to outperform GPT-4 on some (but not all) problems.
This is really great news for open models as many applications will benefit from smaller, faster, and cheaper fine-tuned models rather than a single large, slow, general-purpose model (Llama-2-7B is something like 2% of the size of GPT-4).
GPT-4 continues to outperform even the fine-tuned 70B model on grade-school math question answering, likely due to the data Llama-2 was trained on (more data for fine-tuning helps here).
https://www.anyscale.com/blog/fine-tuning-llama-2-a-comprehe...
You're welcome.
[0] https://github.com/microsoft/azurechatgpt
[1] https://web.archive.org/web/20230814080150/https://github.co...
If you pay, do you get a Ts&Cs that don't contain any wording like this? Still, even if there was no specific "we own everything" statement there could be pretty much standard statement of "we'll retain data as required for the delivery and improvement of the service" which is essentially the same thing.
So, any company that allows it's employees to use chatgpt for work stuff (writing emails with company secrets etc) is definitely not engaging in "secure and private" use.
Unless there is very clear data ownership, for example, customer owns the data going in and going out. I can't see how it can be any different. The problem (not at all)OpenAI has in delivery such service is that in contrast to open source models I'm told there is a lot of "secret sauce" around the model(not just the model itself). Specifically input/output processing, result scoring and so on.
> OpenAI will not use data submitted by customers via our API to train or improve our models, unless you explicitly decide to share your data with us for this purpose. You can opt-in to share data.
> Any data sent through the API will be retained for abuse and misuse monitoring purposes for a maximum of 30 days, after which it will be deleted (unless otherwise required by law).
Now, I think you can do shady stuff with that wording as well, but I guess you can also get sued if you kept or used an unreasonable percentage of your data longer than when you promised to delete it.
Perhaps more nit-pickinlgy specific, they may be compelled by law (the courts or an agency with enforcement capacity) to maintain evidence if there's a warrant or ongoing lawsuit.
I don't think this is accurate. At least in Norway you can't "just not" keep records required by law - any section in a contract in conflict with current law would simply be invalid?
I think the section just clarifies that Microsoft will comply with laws requiring them to keep data (eg the "anti-terror" laws that might require data retention).
Any law. It just makes explicit that a contract can't supercede laws. Even if it was left out, Microsoft is still subject to laws.
On top, you might argue that Microsoft and Azure are easier to trust than a still rather new AI startup.
What kind of volume were you doing and did you use the API for anything other than your listed use case when applying?
I think what happened is the azure subscription was converted from a (multi year) promotional subsidy/discount to a full pay as you go subscription. No change to sub id. Payment methods OK. Everything else continued working, but openai gpt-4 access stopped the next day.
I’d rather use the Azure version because they promise 12-month sunsets vs OpenAI 6-month sunsets for model versions.
Azure is mostly better for production: the developer experience is awful and the default filtering is more aggressive, but you get dedicated capacity by default which improves latency (something you need to negotiate with OpenAI's sales team for otherwise)
It's what happens in the interface, that is your web chat or API call, which is different per implementation. ChatGPT is an implementation that uses that model and its maker OpenAI wants to keep your history for further training.
But what Azure is doing is taking that model and putting it behind an endpoint specific to your Azure account. Businesses have been interested in gpt, so asking for private endpoints. Amazon is doing the same with Bedrock.
In HN-space, it is at its most abstract, idealistic, etc. At the practical level this services is aimed at... it might mean compliance, or CYA. Less cynically, it might mean something mundane. MSFT's guarantee, a responsive place to report security issues.
So calling the repo "azurechatgpt" is misleading. It should really be "sample-chatgpt-api-frontend" or something of that sort.
I don’t know enough typescript to understand where the front end stops and the backend begins I this code
Less than a day later. The last article I see linking to it was published this morning.
Not sure what happened here, but “404’s at just-announced permalinks” seems to be on the rise lately.
Don’t turn me into a late-onset pedant. Fine. URIs are permanent forever! For all resources! ;)
Oh nice!
“But that is lager”
I don't remember seeing this disclaimer on the ChatGPT website, gee maybe OpenAI should add this so folks stop using it.
If you understand what happens on a technical level, it might be possible, but OpenAI has never said this was a risk by using their product.
So in general enterprises cannot allow internal users to paste private code into ChatGPT, for example.
AIUI they are using current chat data for training GPT-5, not re-finetuning the existing models.
Do we know that it wouldn’t have varied in its answer by just as much, if you had tried in a new session at the same time?
Now it seems to know this mathematical property from first prompt though.
Edit: yes
[0]: https://www.theverge.com/2023/3/21/23649806/chatgpt-chat-his...
I see the use of general purpose LLMs like ChatGPT, but smaller fine tuned models will probably end up being more useful for deployed applications in most companies. Off topic, but I was experimenting with LLongMA-2-7b-16K today, running it very inexpensively in the cloud, and given about 12K of context text it really performed well. This is an easy model to deploy. 7B parameter models can be useful.
OpenAI will likely target private consumers while Microsoft focuses on enterprise. I can use my own organisation as an example. We’re an investment bank that does green energy within the EU. We would absolutely use GPT if it was legal, but it isn’t, and it likely never will be considering their finance model is partly to steal as much data as they can. Even if it’s not so polite to say that. This is where Microsoft comes into the picture. In non-tech enterprise you’re buying Microsoft products because everyone wants windows, outlook and office. We can wish it wasn’t like that, but where is the realistic alternative? I’m not anti Microsoft by the way, in all my decades in the enterprise business they’ve easily been the best and most consistent business partner for any IT. When Amazon saw how much money there was on the operations side of EU enterprise they quickly caught up, but Amazon doesn’t sell a Office365 product. So anyway, once you have Office365, you’re also likely to use Teams as your communications platform (which is why there is an anti-trust case against it), Sharepoint as your document platform, and, well, Azure as your cloud platform. Except you might use AWS because Amazon is also great. In some ways they are even more compliant with EU legislation than Microsoft.
But if Microsoft can throw GPT products into Azure the same way they put Teams and Sharepoint into Office365… well, then where is their competition? And having GPT features within Office365 will only further their advantage on the office platform. I mean, there are companies which won’t use Outlook, but there won’t be when ChatGPT writes your e-mails.
So this isn’t necessarily for you. It’s just part of Microsoft’s over all strategy for total IT domination in Enterprise. I mean, we’re going into RPA (robot process automation) a journey I went through in another Enterprise organisation a few years back. Back then you had to consider what go buy, would it be BluePrism, UIPath, automation anywhere, something else? Today there is no competition to Microsoft’s PowerAutomate if you’re already a Microsoft customer. It’s literally $500 a month vs $50k a month… I mean… that’s the future for GPT on Azure.
It’s probably necessary too. Their prices have made a lot of organisations look outside of Azure. Toward places like Hetzner or even self-hosting, but if Azure comes with GPT… well then.
> Starting on March 1, 2023, we are making two changes to our data usage and retention policies:
> OpenAI will not use data submitted by customers via our API to train or improve our models, unless you explicitly decide to share your data with us for this purpose. You can opt-in to share data.
> Any data sent through the API will be retained for abuse and misuse monitoring purposes for a maximum of 30 days, after which it will be deleted (unless otherwise required by law).
Can't you just use the Azure service now?
If I had the time I'd like to play with an MoE of Llama2, as a compare and contrast, but that ain't gonna happen anytime soon.
I've been building https://gasbyai.com, a beautiful chat UI that support self-hosted, with ChatGPT plugins, extract content from pdf/url. GasbyAI supports Azure, OpenAI, and custom API endpoints in case you want to run with your own models
Imo if you're making an open ended chat interface for a business, you're doing it wrong.
With Azure's move to try to internalize any enterprise integration for AI it makes sense to make a chatbot wrapper because its a no-moat move. I think a lot of the "moat" if one can exist in the "chat with your docs" vertical is just integrations into flows and data sources SMB/Enterprises are already using.
For businesses, in my experience, the on-prem thing has been the first decision point - without question. Azure wrapper could be nice to have for those who cannot use chatGPT on the work comp but have access to this instead.
I wonder what kind of hypervisor view it gives to Azure admins for those who use it - it any. Multi-tenant instances was the second highest demand from SMB/Enterprise customers for AnythingLLM.
I'm not familiar with Azure platform.
Is the inference processed on private instance ? I can't imagine how it could be feasible given the hardware required to run gpt3.5/4.
So the best case scenario is:
1. A web ui runs on a private instances. So any user input (chat or files) are only seen by these instances 2. Any chat historisation or RAG is also done on these instances too. 3. Embeddings compuation may possibly be done on the private instance 4. The embeddings are then sent to the Microsoft GPU farm for inference.
So at one point my data has to leave my private network.
The problem is that the data can easily be retro-engineered from the embeddings.
How can this be presented as a private LLM ?
Apparently, many organizations have their own Azure OpenAI deployment and won’t let their employees use the public OpenAI service.
My understanding is that Azure makes sure all network traffic is isolated to their network so they have more controls over how their organization use ChatGPT.
I created a super simple step-by-step guide on how to obtain an Azure OpenAI endpoint & key here:
https://pdfpals.com/help/how-to-generate-azure-openai-api-ke...
Hope it would be useful to someone just getting started with Azure.
[0]: https://boltai.com
[1]: https://pdfpals.com
All I can see is the same product but offered by a larger organization. I.e. they're more likely to get the security details right, and you can potentially win more in a lawsuit should things go bad.
Eventually got more from Open AI so we load balance both. The only difference is the 3.5 turbo model on Azure is outdated.
anybody know why ?
And you can deploy a chat bot from within the Azure playground which runs on another codebase.
Now that Microsoft has an official "enterprise" version out, the floodgates are open. They stand to make a killing.
https://azure.microsoft.com/en-us/explore/trusted-cloud/priv...
https://azure.microsoft.com/en-us/blog/3-reasons-why-azure-s...
I guess I would trust them, since they're big and they make these promises and other big companies use them.
I'd expect this trend of managed ChatGPT clones to continue. You can own the stack end to end, and even swap out OpenAI for a different LLM (or your own model trained on internal company data) fairly easily.
Edit: Uses Azure OpenAI as the backend
Most companies use cloud already for their data, processing, etc. and aren’t running anything major locally, let alone ML models, this is putting trust in the cloud they already use.
EDIT: what you say about existing cloud customers being able to extend their trust to this new thing makes sense, thanks.
Here are a few:
Data privacy
Ownership of IP
Control over ops
The table in the blog lists the top 10 reasons why companies do this based on about 50 customer interviews.
> Private: Built-in guarantees around the privacy of your data and fully isolated from those operated by OpenAI.
Do tell.
It's only going to get more impossible. All that VC money going in at 100x revenue needs a return and they aren't going leave money on the table with full-featured open-source or CentOS type alternatives.
All those data engineering startups, database providers with 'open-source' + cloud hosting, the 'open-source' is going to be just 'open' enough to claim there is some fallback for someone else to pick up the mantle using the community version, if the cloud version gets enshittified beyond reason.
You're not going to even be able to run the full-featured software version on-prem because the economics of cloud are so much better.
Unless you are writing and compiling your own code you are going to be out of luck if your privacy standard is that high. That war has been lost. And Web3 sure ain't gonna save you either.
If we accept this just we accepted the very flawed solutions we were given by corporation regarding social networking and ads, we are going to be stuck with it, suffer the consequences, and there will be no incentive to develop alternatives that actually address issues and work.
Homomorphic Encryption works. It just doesn't work very efficiently right now but that is an intellectual problem that can be solved if we push for actual privacy for this critical technology as it will be fully enmeshed in all parts of our lives.
"Think of the children" if that helps.
Pretty bold thing to say to your potential clients. "You can always tell your employees not to use our product, but they won't listen to you."
I'm wondering as a hobbyist / tinkerer if a solution like this is "affordable" (I know it's all relative)
"I'm not concerned that artificial intelligence will take over the world. I'm concerned that human intelligence has yet to do so."
And the reason is, it's enough for OpenAI to "say" that they're "not going to use your data" - you need a cloud deployment where you can control network boundaries to _prove_ that your data isn't going anywhere it isn't supposed to.
Microsoft is an investor in OpenAI, but does not own it, and they are legally separate companies. OpenAI is not Microsoft and it is factually incorrect to claim that OpenAI is Microsoft.
[1] https://blogs.microsoft.com/blog/2023/01/23/microsoftandopen...
It's not just a straight trade of dollars for shares, but many further contractual obligations.
Thus, it could very well be OpenAI has taken dollars, is commercially selling its technology to Microsoft on terms which aren't special, and sama and the OpenAI executive team and board has independently concluded that engaging in the partnership is a stellar way to grow their OpenAI brand, business and valuation?
we were looking to explore Llama2 for internal use
You can of course run Llama2 in Azure, but you can't host OpenAI models in AWS
Install in 10 minutes.
Make sure you have enough GPU memory to fit your llama model if you want good perf
It means you run the front end (the chat-gui) and the backend code from the repo. This code connects to cosmo-db for uploading documents used for "chat with you pdf" and connects to an OpenAI instance on Azure for the chat inferrence.
One deployment = a deployed model which you can query
On top of that, depending on the model you're using, you also see a cost increment for each 1000 request you make.
Really surprised to see this aggressive of language 1) written down 2) on Github. I'd be pretty pissed if I was OpenAI, regardless of the $10B.
That's how it's been working for months, and if OpenAI objected they would have done something about it.
https://azure.microsoft.com/en-us/pricing/details/cognitive-...
disclaimer/source: I work at Microsoft on Azure/OpenAI
https://www.schneier.com/blog/archives/2023/08/microsoft-sig...
I want to chat and ask about an entire body of knowledge - wiki pages, git commit diffs/messages, jira tasks.
In this day and age, exactly how private does anyone expect their comms/thoughts/files/data to be? I recall reading a recent MS EULA and it seems I have to say three Hail Marys every third Tuesday for using Arch Linux on my PCs. I could install Edge, and did but I don't like the nasty homepage - a bit right wing ... - why on earth is a browser pushing "news"? Its a browser. To be fair I had to dump all the homepage crap that Firefox pushed when I finally dumped anything to do with Chrome.
Please don't use the words private and secure when you have your fingers crossed behind your back.
At the enterprise level where this is intended to be ran, things are much diffrerent.
If you're not aware of the differences or use cases, perhaps you're not the target audience who should be using or configuring it.
Win 10 and 11 are steering you to cloud first, out of the box. That's fine if you like it, but I don't and quite a lot of my customers don't.
The real problem is about data sovereignty. I'm a Brit and ... MS isn't.
You can make the argument for their consumer editions sure, but that's a different product with different features, different price point for different users.
I was gobsmacked to hear a friend say that their work guidance is to use ChatGPT to write letters to external clients for example. I know for sure I'd be insulted if someone sent me paragraphs of text to read created from a sentence long prompt. I'd rather have the prompt, my time is valuable as well.
Cue someone making some horrible error because some crucial information didn't survive ChatGPT->ChatGPT round-trip
Very soon everyone will in effect "hide" behind an agent that will take all kinds of decisions on one's behalf. Everything from writing e-mails to proposals but also to sue someone, make financial decisions, and be a filter that transforms everything going in or out.
I can't imagine this world really. How the hell are people going to compete or stand out? Doesn't it seem that what little meritocracy existed wills soon drown in noise?
The realization that individuals will also have this barrier to the world is even scarier.
If it goes that way we could be looking at a change to society on the level of social media, again. Mad.
I honestly don’t know why we’re so obsessed with having LLMs generate crap. Especially when they’re very capable of reducing, simplifying. Imagine penetrating legal texts, political bills, obtuse technical writing, academic papers and making sense of those quickly. Much more useful imo.
I'm not anti-AI; I've recommended that we use it at work a few times where it made sense and was backed by evidence/bencharmks. But for essentially any problem that comes up someone will try to solve it with ChatGPT, even if it demonstrably can't do the job. And these are not business folks, these are engineering leaders who absolutely have the capability to understand this technology.
We've found some early success selling to companies with older "long-tail" ERP's. I've been finding a new one every day.
Naively asking a chatbot to write for you does not help with this at all.
It would be interesting to try to prompt ChatGPT to ask questions to try to figure out what the user is trying to write and then to write it.
From Microsoft?
Ha.