5,377 karma · joined October 14, 2008
No Bullshit Guide to Mathematics => https://noBSmath.com (high school math review)
No Bullshit Guide to Math & Physics => https://minireference.com (mechanics and calculus)
No Bullshit Guide to Linear Algebra => http://gum.co/noBSLA
No Bullshit Guide to Statistics => http://gum.co/noBSstats (data, probability, and statistical inference)
Elsewhere on social: https://x.com/ivansavz - https://mathstodon.xyz/@ivansavz - https://bsky.app/profile/ivansavz.bsky.social
It's still complicated but the visualizations help a lot!
So, I'm not sure if it's a question of time at all: if a LLM text contains some piece of information beyond the information that went into the prompt, where does this "extra" information come from? [Note, I'm not thinking about facts which could trivially come from the training corpus, I'm thinking specifically as information in the sense of intended message from sender (author) to receiver (reader)]
[1] cf. this comment where I explain this analogy between LLMs and noise channel in communication theory: https://news.ycombinator.com/item?id=49510244
My guess would be "Write a 2000 words blog post based on the analogy between the difficulty of understanding the laws of physics and minified source code."
It's a cool analogy, no doubt, but reading the 2000 words of the article, I felt there was no substance or flesh around it. Okay there were some examples vaguely related to the analogy, and writing was smoothish, but what was the message in the end? What were the new ideas (beyond the analogy)?
Every time I read a LLM generated long form text I am reminded of the communication engineering: communication channels can only decrease the information content. For example, in the communication setup:
Alice sends msg --encode--> codeword --channel--> codeword' --decode--> msg' Bob received
The best we can hope for is that the message was transmitted faithfully, so msg' = msg, since the channel noise can only degrade the information content, specifically, the mutual information is a non-increasing function: I(msg;codeword') ≤ I(msg;codeword),
due to the monotonicity of information (a.k.a. data processing inequality).Similarly, if Alice uses LLM to communicate to Bob, we have:
Alice's idea --encode--> prompt --LLM--> generated text --decode--> idea'
Monotonicity of information tells us I(idea;generated text) ≤ I(idea;prompt).
So I really don't get it, why do people believe that adding LLM generation step could improve the communication of their ideas? It seems directly analogous to passing a message through an additional noisy channel. Just show us the prompt.This book is for medical professionals, but if you find your doctor doesn't have the relevant training then worth to educate yourself. The TL;DR is that you need to go really slowly in the final steps.
Horowitz also has many talks and podcast appearances on youtube. Example old video https://www.youtube.com/watch?v=LRr09mvMEAU but he has many more recent ones.
I don't have any direct experience with this, but I've been watching my best friend struggle with widthdraw symtoms for many year, and I also know from doctor friends that they are not trained on de-prescribing protocols.
I'd like to work with a corpus offline (internal university research data) and I'm hoping I can get everything done without the data leaving the premises.
I guess the biggest bottleneck is going to be for the context window size which won't be able to fit too many result "hits."
Any info or advice would be appreciated.
This is also the strategy I use for editing drafts of my books. I bring a printed draft to someplace nice (e.g. coffee shop or park) and read it all carefully, then I transfer the edits back to the .tex sources. I do several passes of this, until I feel the text + explanations are solid.
Reading on screen just isn't the same...
I remember from my tutoring days how useful they were to organize the different concepts covered in each lesson: I would start with a blank sheet and make the student add concepts to it as the lesson progressed, then by the end of the lesson use the concept map to review what we learned. Specifically, I would ask them to explain in their own words each "arrow" which was a great way to uncover misconceptions and solidify the material.
Once I'm done with editing the current book[1], I hope to have time to work on making dynamic concept maps that you can click on and explore/zoom-in on. I feel it would be cool to jump between detailed view (concepts), intermediate scale (topics), and high-level view (subjects).
and the printable concept maps here: https://minireference.com/static/conceptmaps/linear_algebra_...
Perhaps there might be some "metric" that can be calculated automatically from git commit history, but it would be nicer if the author can manually allocate "weight" to each passage of the text, based on what they perceive is most important (read the dark regions for the the deepest insights).
--
Use Case 1: Emphasis in office communications. When employee A communicates with employee B, they can use the text background to communicate the energy invested by employee A to produce the email/report/memo/doc in question.
Example 1: Developer spent 1 month rewriting the login system to a clean, feature-identical, drop-in replacement Auth API. When he announces to the CTO, the text is in solid grey, communicating the number of hours that went into this project and hence it's importance (better read all the details). dark highlight = I invested a lot of energy in this; I know you're a business person now, but I need you to read this and act on it!
(current alternative way to "highlight" actions is to link them to financial benefits for the company, e.g. the new Auth API will save us hours of dev time and allow us to integrate with platforms X, Y, Z, but I think the "number of hours I put into this" would be a good metric to show as well. It's proximal to the developer's emotional investment in the thing)
--
Use Case 2: Emphasize key ideas in educational text. Educators can communicate the relative importance of different parts of the text. Instead of "energy that went into creating this text," the more useful signal would be to tell readers how much mental energy it will take them to understand a given concept. A sentence that explains some key concept can be shown in dark background to say "this is deep" and also "this is worth learning about specifically within this section." Example 2: An author explains a sequence of 10 complicated steps that are part of some process, and indicates by text background color which steps are important and which are just technical minutiae.
* I write math and science textbooks, and I'd love to have such a sidechannel to the reader to communicate "importance" ... I hack by marking some key stences as bold, but it could be better.
--
Use Case 3: Ethical labelling of genAI output. It is now commonplace within the academic world to require disclosure of genAI use when creating scientific publications and other docs. This is a good practice, but it leaves too much degrees of freedom to the author about the level of disclosure they make. A word-for-word provenance metadata "channel" in parallel with the final text of the document could be a very useful thing to have (in an ideal world). In the real world, few academic would admit to heavily leaning on genAI, but at least we can shame them for not providing the "detailed provenance" track along with their text.
In contrast, people who are using genAI unashamedly (half the population) could prove how un-ashamed of their genAI usage they are by specifying "All genAI" in the "provenance channel" for all the posts they publish. If you like the SLOP or you think you can control the SLOP, I won't judge you, but please let me know so I read the text differently...
Example 3: Author A asks editor E to review a book draft. The manuscript clearly shows which parts of the book were written by the author and which parts are genAI. The editor knows which parts to focus their attention on (what the author is saying), and which parts are just filler. For bonus points, the author could also disclose the harness+context+prompt they used to produce the text (like a view source affordance that comes with any genAI text passage).
I also have some notebooks with SymPy code examples here: https://github.com/minireference/noBSLAnotebooks
I'm conceptualizing a piece of knowledge as an interface that can be `implemented` but with different classes (explanations renderings for different audiences).
For example, the "derivative interface" represents knowledge of the concept of derivative operations and basic skills to compute derivatives of various functions. The interface doesn't specify HOW to teach this topic or HOW DEEP, so there are multiple implementations:
- basic visual explanations (for kids)
- basic algebra steps (for high school)
- standard explanation (for undergraduate students)
- compact explanation (a reviee for grad students)
The above implementation are polymorphism due to the "reader level of knowledge," but there could be other, e.g. derivatives explained using code like in Sec 4.1 in this calculus tutorial[1].It would be A LOT of work to produce all these explanations but it would make for a kick ass math textbook that you can pick up and learn, no matter what your level is (instead of getting lost or bored and looking for another resource).
[1] https://minireference.com/static/tutorials/calculus_tutorial...
I found this article[1] by the author[2] that explain their motivation for the book. It has lots of deep insights about the textbook publishing industry and definitely worth a read.
[1] https://framablog.org/2022/01/20/mais-ou-sont-les-livres-uni... and auto-translation to EN for ppl who don't speak French: https://framablog-org.translate.goog/2022/01/20/mais-ou-sont...
For anyone interested in checking out the book, there is a PDF preview here[1] and printable concept maps[2], which should be useful no matter which book you're reading.
[1] https://minireference.com/static/excerpts/noBSmathphys_v5_pr...
[2] https://minireference.com/static/conceptmaps/math_and_physic...
I was working on web copy describing how crazy the mainstream textbook prices are, and used the price C$300 for the calculus book, trying to be flippant (to exaggerate the competitor price to make my prices look better). I decided to check the price in the bookstore, and to my surprise the price was even higher than that! (sold as bundle: book + exercise manual + solutions manual). When your real prices are higher than the pricing people use as hyperbole, you know there is a problem.
It makes no sense—for a subject that has been around for 300+ years, and virtually unchanged for the past 100.
It all depends on the integrity of the researcher, which in turn depends on their upbringing (the example set by their academic advisor).
Relevant links to two podcast episodes about p-hacking:
https://web.archive.org/web/20260314130229/http://silas.psfc...
One thing that works very well for me (in a different context) is to ask to return two lists:
- Things that I must absolutely fix (bugs, typos, logic mistakes, etc.)
- Lesser fixes and other stylistic improvements
Then I look only at the first list.
https://fr.wikisource.org/wiki/Trait%C3%A9_des_excitants_mod...
The part about coffee is halfway down the page under the heading §III — du café.
Can you say a few more words about the library https://github.com/standardebooks/tools ? Can it generate ePub3 from markdown files or do I have to feed it HTML already. Any repo with usage examples of the `--white-label` option would be nice.
Scrolling is no longer interesting, and food looks un-appetizing. Making the digital reality look boring is a good deal to make the real world look more exciting.
Thanks to comments from @jtbaker and @SkyPuncher I just added a shortcut to the "pull out" menu so I can now turn off when I need to work with pictures where colors are important.
The website has all the notebooks from the book, as well as well as the complete tutorials on the tech stack (Python, Pandas, Seaborn).
For everyone interested, check out the extended preview PDFs:
- Part 1: DATA and PROBABILITY https://minireference.com/static/excerpts/noBSstats_part1_pr...
- Part 2: STATISTICAL INFERENCE https://minireference.com/static/excerpts/noBSstats_part2_pr...
(Y @ X)[None]
# array([[14, 32, 50]])
but `(Y @ X)[None].T` works as you described: (Y @ X)[None].T
# array([[14],
# [32],
# [50]])
I don't know either RE supposed to or not, though I know np.newaxis is an alias for None.