Show HN: DocsGPT, open-source documentation assistant, fully aware of libraries
github.com
github.com
It also appears the Github library for DocsGPT was created shortly after the release of KnowledgeGPT.
I thought about the same problem. Companies have lots of marketing material and data sheets that are not easily queried but would be very useful to support staff when dealing with customer's questions.
I didn't build it, they did. They get my support.
Productivity tools shouldn't depend on external parties like that.
If not OpenAI, Google will sell general use AI APIs
But the problem will be solved regardless of GPT.
I don't think the whole stack relies on GPT for this solution.
Some parts maybe, but the whole of the project is more than that.
Hey, I don't agree with "... made with Rust/Go/React/X" slogan too. It is just pointless.
Thanks so much!
Caveat emptor: I didn't try this out yet. No idea whether this would work.
[the repo](https://github.com/karfly/chatgpt_telegram_bot)
The KnowledgeGPT repo linked by another commenter seems more interesting.
Kudos to you, and I hope you keep that excitement going!
>"This will return a new DataFrame with all the columns from both tables, and only the rows that match the 'key' column".
That is incorrect. The 'how' parameter is 'left' not 'inner'.
That this is not a problem for many is worrying.
I will be much more excited when AI can explain undocumented systems to me. This feels like it can't be far away, and it will be a game changer.
I guess for this to be helpful it is gonna need out-of-band info, but maybe just the git log would be a pretty good start. If you could add a mailing list or chat history of developers I imagine things could get more powerful.
It's like complaining that a musician cannot practice their instrument without software that requires you to be always-online.
> Even if it was open source no user would be able to run it locally anyway.
You are just stating this, it does not make it true. Several of us are running GPT-3 workloads locally.
> Several of us are running GPT-3 workloads locally.
Several people out of 8 billion is the same as no-one. The point made was that anyone should be able to run this document assistant on their local machine.
Unless you work for OpenAI (and your “local” machine somehow has 1TB+ of VRAM, equivalent to roughly 25 A100s), this cannot possibly be true…
apart from that, several pre trained corpuses had been around for a while
Do you mind sharing how? I didn't think the model was available to download and thought it was api only?
When I was in school, in AI/ML classes we needed to build our own models. So now I'm guessing that you use existing models...
I'm trying to gauge how behind the times I am.
>> How do the responses compare to auto-summarization in terms of Big E notation and usefulness?
> Automatic summarization: https://en.wikipedia.org/wiki/Automatic_summarization
> "Automatic summarization" GH topic: https://github.com/topics/automatic-summarization
Though now archived,
> Microsoft/nlp-recipes lists current NLP tasks that would be helpful for a docs bot: https://github.com/microsoft/nlp-recipes#content
NLP Tasks: Text Classification, Named Entity Recognition, Text Summarization, Entailment, Question Answering, Sentence Similarity, Embeddings, Sentiment Analysis, Model Explainability, and Auto-Annotatiom
Are people asking a question under the hood? Like “explain these 30 code files to me”
Using Ada or Babbage is about 1% of the cost of Davinci (and Curie is 10% the cost of Davinci).
Without any real tuning, this responds quite promptly (and the various tests I've done, correctly):
curl https://api.openai.com/v1/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"model": "text-ada-001",
"prompt": "Identify if the following question is about the Olympic Games. Answer with {yes}, {no}, {maybe}.\n\nWhat category is pole vaulting in?",
"temperature": 0.7,
"max_tokens": 38,
"top_p": 1,
"frequency_penalty": 0,
"presence_penalty": 0
}'>> Yes
Well, at least you can't say the documentation is out of date.
It seems like adding a translation layer into conversational speech could make good docs harder to use?
And if theyre not good docs to start, what will gpt base its answer on?
Could be useful for a question that spans multiple technologies maybe
> You can use the terraform-aws-api-gateway module. This module provides a set of Terraform resources for creating and managing an API gateway. It allows you to define the API gateway, its resources, methods, and stages. It also provides support for custom domain names, API keys, and usage plans.
I followed up with "Show me the hcl configuration for it."
> There is no hcl configuration for it.
Would this support Markdown format? Context, I'm using Hugo with Markdown + injected code snippets. Might need to actually crawl the site...
Do you think including Slack Q&As in training would help too?
https://github.com/arc53/docsgpt/blob/main/scripts/ingest_rs...
This is the prompt that is used: https://github.com/arc53/docsgpt/blob/main/application/combi...
And this is where it calls the openai. Looks pretty straightforward. https://github.com/arc53/docsgpt/blob/main/application/app.p...
Really sad but but it seems like most people seem to not care about privacy even for the most sensitive parts of their lives.
Or even worse they mistrust "the government" but place great trust in "companies which are easily pressured or raided by more than one government".
It doesn't make sense at all.
(Edit: I don't think a well run government in a functioning democracy is inherently evil but pretending that companies are not collaborating because they are "private entities" is foolish)
Control is why I want to see more "open source" version of machine learning models be it over data or auditing the output and if I can p2p download a model I know we're in a good place.
Take the internet, I could see my path being very different if we had to pay per minute like a phone plan to get information, having open access was a big bootstrap here.
Finally of course there are problems with the Linux project, but would servers dictated only by M$oft/0r@cle be a positive outcome, gonna say no.