Snomed CT Entity Linking Challenge
drivendata.org
drivendata.org
Software Development IDEs have (optional) autocomplete;
Unstructured Medical Coding interfaces could also have autocomplete,
such that when you type `icd:` or `snomed:` it presents a search interface for that particular medical terminology / vocabulary / system of classification / categories.
GNU Health > Issue tracker > "Freetext ICD-10 references as URIs (e.g. icd10:A01)" (2013) https://lists.gnu.org/archive/html/health-dev/2013-12/msg000...
To prompt them at that time for whether or not this referenced thing in a dictated .txt file is actually a thing with a URI and have them confirm that those are the correct annotations for their input would save a lot of time and money.
Requisite study for Linked Data clinical coding / informatics:
- HIPAA and unstructured notes, HIPAA and Linked Data; Informed Consent; Precision Medicine
- Tokens; NLP, stemming, conceptual entity recognition
- The LODcloud; a great big graph of Linked Open Data (that our data does not yet link to, is siloed separately from, does not yet have references to existing URIs in)
- RDFa: RDF-in-html-Attributes
- JSON-LD: JSON Linked Data
- Schema.org/MedicalEntity ; /docs/schemas.html -> "Health and medical types" https://schema.org/docs/meddocs.html
- TTS Text-to-Speech, Speech-to-Text; Multimodal models with RAG, LORA, RLHF, Transformer Self Attention tensor Networks, Benchmarking OpenAI Whisper with relative performance metrics, ONNX
- FHIR for EHR data portability (JSON-LD,)
UX Improvements:
- More IDE-like unstructured data entry
- Auto-annotate this text field in place
- Auto-annotate this pasted text
- Auto-annotate this text field and display the linked data annotations
- Display the revisions of the unstructured note -> annotated linked data document
- Autocomplete with search from e.g. icd:
- Frequently used annotations per Clinician/Provider
- Recently used annotations per Clinician/Provider
- Hover over visually distinguished auto-recognized entities and approve/clarify/reject
- Indicate which prov:Agent added which annotations; human (name(s)), AI (signed git revision of repo URI)
- Indicate degree of confidence in annotation (note that AGI hypergraph systems have TruthValue and also AttentionValue, like attention networks; and that AGI systems may someday also support CDSS: Clinical Decision Support Systems and Health Decision Support Systems)
Most EHRs already have features for building custom data entry template forms for common situations, which can allow for selecting specific codes. Large provider organizations typically have internal committees to design these forms in order to ensure some level of consistency in patient charts.
Transmission chain method > Application in studies of memory: https://en.wikipedia.org/wiki/Transmission_chain_method
https://en.wikipedia.org/wiki/False_memory_syndrome#Evidence... :
> Human memory is created and highly suggestible
Though nowhere is it shown that clinicians are even competent enough to add RDFa linked data annotations to their TXT unstructured notes with an encabulator.
It's clear that there's value in disambiguation; did they say "biomemetic" or "liquid metal" days ago on a dictaphone in a loud environment?
You can add Linked Data (JSON-LD RDF) annotations with threaded markdown to document Trump vectors with W3C Web Annotations.
This [SNMOED-CT, ICD10-] competition could be scored with Web Annotations as the schema for the output.
W3C Web Annotations open source implementations:
- hypothesis/h: https://github.com/hypothesis/h
ElasticSearch's Term Vector optionally returns with_positions_offsets, with_positions_offsets_payloads: https://www.elastic.co/guide/en/elasticsearch/reference/curr...
Meilisearch has showMatchesPosition: https://www.meilisearch.com/docs/reference/api/search#show-m... :
> `showMatchesPosition` returns the location of matched query terms within all attributes, even attributes that are not set as searchableAttributes.
Meilisearch > Comparison to Alternatives; Algolia, ElasticSearch, Meilisearch, Typesense: https://www.meilisearch.com/docs/learn/what_is_meilisearch/c...
And then now instead, Vector search engines: TODO
NER: Named Entity Recognition: https://en.wikipedia.org/wiki/Named-entity_recognition
awsome-medical-coding-nlp: https://github.com/acadTags/Awesome-medical-coding-NLP
awesome-ehr-deep-learning: https://github.com/hurcy/awesome-ehr-deeplearning
awesome-ner: https://github.com/smiyawaki0820/awesome-ner
awesome-bioie > Research groups: https://github.com/caufieldjh/awesome-bioie#groups-active-in...
SNOMED-CT as RDF: https://sphn-semantic-framework.readthedocs.io/en/latest/ext...
...
SNOMED-CT is a Medical Terminology.
SNOMED-CT: https://en.wikipedia.org/wiki/SNOMED_CT
Traditionally, a Terminology Service provides a query endpoint over a schema like SNOMED-CT (which is a nonredistributable medical XML schema, which is challenging for Open Linked Data):
https://github.com/NCIP/lexevs
https://github.com/OpenConceptLab
https://openconceptlab.org/terminology-service/
https://github.com/MedevaKnowledgeSystems/pymedtermino/blob/... :
> For SNOMED CT, ICD10 and MedDRA, the data are not included (because they are not freely redistribuable) but they can be downloaded in XML format. PyMedTermino includes scripts for exporting these data into SQLite3 databases.