HNHacker News
TopNewBestAskShowJobs

dsubburam

808 karma · joined October 23, 2014

CEO at https://www.chatlaw.us Writer at https://deepsub.substack.com/about
submissionscomments
dsubburam··on Show HN: Q&A with AI Trained on Bankruptcy Law
I can think of a way for it to just decline to answer, after noticing that the question is not relevant to its target area. Thanks for the feedback!

There's a Reddit forum[1] for discussing personal bankruptcy, btw, if you're curious what the process looks like. (I found it interesting myself.)

[1] https://www.reddit.com/r/Bankruptcy/

dsubburam··on Show HN: Q&A with AI Trained on Bankruptcy Law
How much do you think I ought to reveal? There are well-funded startups like harvey.ai[1] in this space that I want to stay ahead of.

[1] https://www.harvey.ai/blog See second story announcing their Series A.

dsubburam··on The Alexander Piano
FWIW I've noticed that "rattling" (which I think I hear in the Liszt Funerailles video) in some K-pop videos featuring electronic low bass at the ~40Hz level. I think it might be natural and not an artifact or flaw.
dsubburam··on The Alexander Piano
Conversely, if composers knew of the Alexander piano's existence they may have written specifically for it taking advantage of the bass' clarity.

I just heard the Liszt Funerailles played on the Alexander piano (YouTube on the embedded main page), and yes it sounds very unusually melodic in the bass.

dsubburam··on What we know about LLMs
'Formal Algorithms for Transformers'[1] is a proper account of the architectures and what tasks they naturally lend themselves to, by authors from DeepMind. See sections 3 (Transformers and Typical Tasks) and 6 (Transformer Architectures).

Not much on empirical observations, though.

[1]https://arxiv.org/abs/2207.09238

dsubburam··on Attention Is Off By One
This part of his post where he explains vector embeddings of the input/output tokens just looks wrong to me:

>This vector seems to get taller every model year, for example the recent LLaMA 2 model from Meta uses an embedding vector of length 3,204, which works out to 6KB+ in half-precision floating-point, just to represent one word in the vocabulary, which typically contains 30,000 - 50,000 entries.

>Now if you’re a memory-miserly C programmer like me, you might wonder, why in the world are these AI goobers using 6KB to represent something that ought to take, like 2 bytes tops? If their vocabulary is less than 2^16=65,384, we only need 16 bits to represent an entry, yeah?

>Well, here is what the Transformer is actually doing: it transforms (eh?) that input vector to an output vector of the same size, and that final 6KB output vector needs to encode absolutely everything needed to predict the token after the current one. The job of each layer of the Transformer is quite literally adding information to the original, single-word vector. This is where the residual (née skip) connections come in: all of the attention machinery is just adding supplementary material to that original two bytes’ worth of information, analyzing the larger context to indicate, for instance, that the word pupil is referring to a student, and not to the hole in your eye.

Firstly, he is confusing representation with encoding--he's right that 2 bytes is enough to encode any token. That is in fact approximately how it's done: a code book is indexed into (with a longint in pytorch, at least last I worked with it ~6 months ago). The purpose of the embedding is to allow the model to learn a representation of the token, a la word2vec. (Though this representation is purely based on the characters comprising the token and does not distinguish between "student" and "eye" in the case of "pupil" as in his example.)

Secondly, his description of each layer's function as adding information to the original vector misses the mark IMO--it is more like the original input is convolved with the weights of the transformer into the output. I am probably missing the mark a bit here as well.

Lastly, his statement that the embedding vector of the final token output needs all the info for the next token is plainly incorrect. The final decoder layer, when predicting the next token, uses all the information from the previous layer's hidden layer, which is the size of the hidden units times the number of tokens so far.

dsubburam··on Harry Frankfurt has died
What had stuck in my mind when I'd read of his work some years back was his idea of "second order desires", wanting to have the desire for X even if you don't have the desire for X (e.g. for healthy food), and how having these second order desires can be seen as a form of free will--without second order desires, you are simply driven by your first order desires (like an animal).

Just googled it for a refresher and found this[1] comprehensive but readable explanation of the above idea. The obituary refers to it briefly as well.

[1] https://philosophy.tamucc.edu/notes/frankfurts-theory

dsubburam··on Aristotle’s Rules for Living Well
Am reading the book[1] slowly, two pages at a time, taking notes.

A recent note I took from Ch 7: Unrestraint and Pleasure

"...virtue keeps the source safe, while vice destroys it, and in actions the source is that for the sake of which one acts ... so neither there nor here is reason able to teach anyone the sources, but here it is virtue, either natural or habituated, that directs one to right opinion about the source."

Reminded me of Buddhist teachings that say that moral conduct is necessary and you can't just think/meditate your way to wisdom/enlightenment[2].

[1] Nicomachean Ethics, Aristotle. tr. Joe Sachs [2] cf. https://puredhamma.net/living-dhamma/transition-to-noble-eig...

dsubburam··on What is a Vector Database? (2021)
Because the model used to compute the embeddings is the same across scenarios. You can infer meaning for each dimension by checking which inputs get embeddings that have large values for the dimension.

If the inputs are images, you may find that some dimension scores e.g. how much blue there is in the image. Though often it's not that simple (there could be multiple dimensions that relate to how blue the image is, especially if the embedding dimensionality is large, which it does tend to be these days. Though you could reduce the embedding dimensionality first using PCA, and see what input images correspond to high/low values of the first principal component, etc.).

dsubburam··on Half of vinyl buyers in the U.S. don’t have a record player: study
> Well, the objective sounds quality is always _worse_ than eg CDs.

Not necessarily true. If you have a good pair of speakers, a good amplifier, but a bad DAC[1], a CD can sound worse than vinyl (whose output does not need to go through a DAC). Old CD players (like mine) have dated built-in DACs, so this is not too exceptional a situation.

The above comment holds even if the vinyl was made from a CD source, since the vinyl maker could have used a quality DAC that's better than your CD player's.

For a long time, I didn't understand why my FM radio channel (WQXR) sounded better than my CDs. Turns out my CD player's DAC was poor in comparison to what the radio station was using to play their CDs.

[1]Digital to Analog Converter.

dsubburam··on Transformers from Scratch (2021)
An early explainer of transformers, which is a quicker read, that I found very useful when they were still new to me, is The Illustrated Transformer[1], by Jay Alammar.

A more recent academic but high-level explanation of transformers, very good for detail on the different flow flavors (e.g. encoder-decoder vs decoder only), is Formal Algorithms for Transformers[2], from DeepMind.

[1] https://jalammar.github.io/illustrated-transformer/ [2] https://arxiv.org/abs/2207.09238

dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
OP here. We now have a Chrome extension[1] to post whichever webpage you're on to klavier.ai for subsequent Q&A. Avoids hassle of having to copy 'n paste the URL.

We are slowly working through the issues reported in this thread. Thanks for the kind and constructive feedback!

[1] https://chrome.google.com/webstore/detail/qa-with-klavier/jb...

dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
BeautifulSoup seems to work well for parsing. For your other question: something like that!
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
4: We would process the entire website, and be able to cite the page/section from where the answer came from, so the customer can look up the original documentation to confirm.

And given your explanation of your usecase, this feature looks more compelling to build out. Would you consider messaging me (email in my profile)? We'd love to chat, and maybe roll out a solution for you as a pilot customer.

dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Thanks. Just hearing about chatpdf.com. Trying them out. Thanks for the tip!
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Not exhaustively, but anecdotally, it's good. We find that few questions end up having to retrieve relevant information from multiple document locations and combine. Maybe a good test here is giving it a HN thread and asking to summarize the positive and negative reactions from commenters?
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
1: We are finding out. Someone else mentioned: https://github.com/whitead/paper-qa We're hoping to keep our service be accessible and easy to use, and add features. Such as from your other questions...

2: We are thinking of the website integration. Do you think OpenAI may release this too? Questions received by email is a new idea that sounds interesting!

3: Thanks for the suggestion – we will look into it.

dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Re: your Edit, it's possible that your questions were follow up questions, which are difficult to make sense of on their own--the service at the moment starts from scratch for each question (has no memory of previous questions and answers).

We'll look into adding memory (either as a default or as an option).

dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
The crash was unrelated to tech and a provider billing issue. Fixed. We do traverse the full document, all 2,000+ pages that we support. Give it a go!
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Noted!
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
DANmode below was right. Now back up!
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Generous enough for most use cases, without hogging our compute and storage resources (currently not at scale).
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Back up. Try again. Sorry!
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Looking into it...
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Yes, the limit is ~2,000 pages; so have at it! As for costs, it's manageable so far but yes, we'd need to figure out a business model. We will likely keep a basic version such as the current one free.
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Appreciate your writing in, and glad you're finding it useful. Would it help if the service shows where in the document the answer was fetched from? (we could work on adding it.)

(And yes, I feel the anxiety too--to keep up with what people are doing with the tech!)

dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
I am on the waitlist for Bing and can't check directly--would it also answer specific questions about the doc? (Rather than summarize.)

Our site is meant for Q&A, and has a layer of tech that finds the sections in the large document that are relevant to the Q first. This will not work well in general for summarization on unstructured content. But most content tends to be structured and in practice we are finding that the approach still works on e.g. news articles, wikipedia articles, blog posts. It's almost as if where it doesn't work, a human would have trouble too. (e.g., on a long rambling HN thread).

dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Thanks for the feedback. It, unlike ChatGPT, doesn't retain question history--so it didn't know what "it" referred to in your final question.

Given your comment we are going to consider retaining question history (or offer an option to do so)!

dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Thanks for the feedback, Joe. We'll look into what's wrong with parsing the React/SPA websites. And maybe do some post-processing for readability. Improved model is in the works too--so hope you give us a try again in a week or so.
dsubburam··on Show HN: Document Q&A with GPT: web, .pdf, .docx, etc.
Whoa. Maybe see your story in Show HN someday, if you'd share!
← PreviousPage 3 of 5Next →