4,868 karma · joined June 18, 2009
Maybe I'm missing how people use AI these days but when I have an agent working locally, all git commits have my authorship attached.
> As rescue divers searched for the boy's body, we deliberated whether to attempt resuscitation and likelihood of meaningful neurologic recovery of a child submerged for at least 90 minutes. We reviewed literature for guidance2-4,6 and drew from institutional experience with a 2-year-old submerged in ice water for 40 minutes who received 101 minutes of CPR.3 The toddler recovered with no sequelae. For our current patient, the decision was made to resuscitate and rewarm the boy because of his young age and protective effects of ice water submersion. We reasoned that if meaningful neurologic function were not observed after rewarming, end-organ preservation on ECMO may allow family goodbyes and organ harvest for transplantation to give other sick children the gift of life.9 This important point should be considered by providers faced with the difficult decision to attempt resuscitation of a patient with asystolic hypothermia >90 minutes.
If you look at Table 3 you can see the difference in performance between for example GPT 5.5 and Opus 4.7 for each of the 20x 100 runs:
- GPT 5.5: 1389/2000 questions answered, of which 1043 were correct (75%)
- Opus: 1306/2000 questions answered, of which 294 were correct (22%)
So while you can claim that Opus solved 40% of the problems it still had a failure rate of 78%. That means if you chose this model to answer your homework question, there is a good chance you would fail.
Perhaps a more useful benchmark for future models is measuring how many of these types of questions they can answer in one shot. I.e. how confident can you be when using them for real world tasks.
Having said that, have a process to automatically grab screenshots is going to make it significantly easier for a developer to update the docs so the motivation to keep the text up to date is going to be much higher.