607 karma · joined January 7, 2018
https://samrawal.com
I’ve heard others say before that real clinical education starts after medical school and once residency starts.
[0] https://samrawal.substack.com/p/the-human-ai-reasoning-shunt
I wrote a little bit more of my thoughts here, in case it’s of interest to anyone: [0]
On that same vein, I recently made a tool I wrote for myself public [1] - it’s a “copilot” for writing medical notes that’s heavily focused on letting the clinician do the clinical reasoning, with the tool exclusively augmenting the flow rather than attempting to replace even a little bit of it.
[0] https://samrawal.substack.com/p/the-human-ai-reasoning-shunt
The buttons are stored to localstorage, as is the editor text - everything stays locally (besides what is sent to the LLM). I'm planning on a simple import/export mechanism for transferring buttons across different browsers/computers!
Reminds me of org-mode in Emacs - overall, I'm really liking the cross-platform sync and Markdown interface out-of-the-box.
>> “you've tried response_format: "json" and function calling and been disappointed by the results”
Can anyone share any examples of disappointments or issues with these techniques? Overall I’ve been pretty happy with JSON mode via OpenAI API so I’m curious to hear about any drawbacks with it.
- Automatically loading/unloading models from memory - just running the Ollama server is a relatively small footprint; every time a particular model is called it is loaded into memory, and then unloaded after 5 mins of no further usage. It makes it very convenient to spin up different models for different use-cases without having to worry about memory management or manually shutting down those tools when not in use.
- OpenAI API compatibility - I run Ollama on a headless machine that has better hardware and connect via SSH port forwarding from my laptop, and with a 1 line change I can reroute any scripts on my laptop from GPT to Llama-3 (or anything else).
Overall, at least for tinkering with multiple local models and building small, personal tools, I've found the utility:maintenance ratio of Ollama to be very positive -- thanks to the team for building something so valuable! :)
It took me a while to buy in to high-volume memorization as a learning technique (especially coming from CS, where memorizing facts is not a huge emphasis). After a while though, I started recognizing how the quick recall encouraged by the system enhanced my understanding of concepts vs replacing it (I wrote about this a couple years ago [0]).
[0] https://samrawal.substack.com/p/on-the-relationship-between-...
After using it for a while though, I began to value how the quick recall encouraged by the system actually seemed to /enhance/ my deeper understanding of concepts, rather than replace it (I wrote a short post about this a couple years ago [0]).
[0] https://samrawal.substack.com/p/on-the-relationship-between-...
“Developed by Tencent's ARC Lab, LLaMA-Pro is an 8.3 billion parameter model. It's an expansion of LLaMA2-7B, further trained on code and math corpora totaling 80 billion tokens.”
I often hear something in a podcast that I want to write down, but don't want to peck out notes on my phone, so I built a GPT Vision-powered tool to help.
Workflow:
- hear something interesting → just screenshot lock screen
- use screenshot to auto-identify podcast, episode, and timestamp
- summarized bullet points at that timestamp via GPT, logged to a virtual notebook
[0] some screenshots - https://twitter.com/samarthrawal/status/1735133820008141170
[1] demo vid - https://twitter.com/samarthrawal/status/1735852196493971593
[0] https://github.com/samrawal/chatgpt-localfiles
[1] https://samrawal.substack.com/p/example-driven-development
[0] Training Llama-2-chat: Llama 2 is pretrained using publicly available online data. An initial version of Llama-2-chat is then created through the use of supervised fine-tuning. Next, Llama-2-chat is iteratively refined using Reinforcement Learning from Human Feedback (RLHF), which includes rejection sampling and proximal policy optimization (PPO).
alias op='open .'>> Total mass of 3,291 kg (GPUs only; not counting chassis, system, rack)
The side project ended up being a single-file HTML file I could throw into a browser tab, so I decided to put it up and share on GitHub :)
>> Now - throw a punch of clinical guidelines in a vector database and give it context
I built MedQA (https://labs.cactiml.com/medqa) as a way to explore how GPT could be used through providing clinical guidelines as context + pretty extensive prompt engineering. There are definitely limitations and the accuracy/constraining "hallucinations" has to meet a very high bar, but I've found it interesting-to-helpful several times while on rounds at the hospital as a med student.
Some functionality made possible through GPT that I am excited to explore further:
- Interacting with the guidelines in a question-and-followup format: https://twitter.com/samarthrawal/status/1620547117717786624
- Answering with different levels of complexity depending on user: https://twitter.com/samarthrawal/status/1636137740390498306
- Proactively detecting if the answer lends itself well to a table format, and optionally generating a table to address user query: https://twitter.com/samarthrawal/status/1631759290414276615