ICD-10 Diagnosis Extractor  (version 3.0)
=========================================

NOT MEDICAL ADVICE. NOT LEGAL ADVICE. NO TOOL CATCHES EVERYTHING.
  The first time you open the window you must read the terms and tick the box
  before the program will run. Your acceptance is recorded in
  icd_tool_accepted.txt in this folder. You are asked again after an update.
  Read the terms any time from  Help > View terms of use.
  Whatever this tool reports, confirm it against the original record yourself.

WHAT IT DOES
  Reads PDF, Word (.docx), Text (.txt), HTML (.htm/.html) and clinical XML
  (.xml - CDA/CCD export) files and pulls out
  every valid ICD-10-CM diagnosis code - and diagnoses named without a code -
  writing a plain-text report:

     file | page/line | code | description | nearest date

  Entries coded under ICD-9-CM or SNOMED CT are listed too, tagged with that
  system so they are not mistaken for ICD-10. A SNOMED tag also says what the
  entry is - a disorder, a finding, a procedure, an event or a situation.

  The report ends with a plain list of just the ICD-10 codes, ready to paste
  into a lookup tool such as  https://ratemyvso.net/dc/icd-codes .
  ICD-9 and SNOMED codes are deliberately left out of that list - that tool maps
  ICD-10, and feeding it other systems gives wrong answers. In the window, the
  "Copy ICD-10 codes" button puts exactly that list on the clipboard.

HOW TO RUN  (Windows)
  Keep icd_extract_gui.exe and icd10_data.tsv together in the same folder.

  The window:
     Double-click  icd_extract_gui.exe .
     Read the terms, tick the box, click Accept (first run only).
     Click "File..." or "Folder...", tick options if you like (e.g. 'Skip scanned pages' = faster but may miss data), press Extract.
     The progress bar moves while it works and the status line beside it shows
     the page it is on. Press Stop to end a long run early - what was found so
     far is kept, and the report is marked as a partial scan.
     When it finishes: "Copy ICD-10 codes", "Open report", "Open folder",
     "Create Claim File Handoff".
     "Clear" resets the window for the next job. Enter also starts a run.
     The folder you used last is remembered for next time.

THE CLAIM FILE HANDOFF (optional)
  After a run finishes, "Create Claim File Handoff" writes a second document
  beside the report, named  <report name>_claim_handoff.txt . It ORGANIZES what
  the run already found. Nothing is rescanned and nothing is added: it is the
  same information as the report, arranged for whoever works on the claim.

  Five sections, in the order you actually need them:

  1. RUN SUMMARY   how many entries were kept and where they went, plus a short
                   name for each source file (F1, F2, ...) so a long filename is
                   not repeated on every line.
  2. POTENTIAL VA DIAGNOSTIC-CODE RESEARCH MATCHES
                   the useful part first. Entries are grouped by the diagnostic
                   code the local reference returned for them. Grouping means
                   the reference gave the same code, NOT that this document says
                   they are the same condition. Every original code and
                   description is kept exactly as the record wrote it.
  3. OTHER DIAGNOSES AND FINDINGS WITH NO LOCAL VA DC MATCH
                   said once for the whole section instead of after every entry.
                   No local result does not mean the entry does not matter.
  4. ADMINISTRATIVE, PROCEDURAL, AND ENTRIES TO VERIFY
                   things that are NOT diagnoses and are no longer presented as
                   ones: Z-code encounters and status codes, SNOMED procedures
                   and events, and named entries the record did not say enough
                   about. An entry that reads like an ordered laboratory test
                   (for example "HGB A1C (Tosoh G8) (83036) Ordered") is put here
                   with the reason printed, never called a diagnosis.
  5. DETAILED SOURCE-PAGE INDEX
                   the complete evidence, page by page, exactly as before.
                   Nothing was removed to shorten the sections above: every
                   entry appears once up there and once down here.

  Reading the compact columns:
     SOURCE   F1 p5/d2 is source file F1, extracted page 5, the page the
              document itself numbers 2. No /d means the document printed no
              page number of its own, and "F1 line 12" means the file has no
              pages at all.
     DATE     a trailing * means APPROXIMATE, the nearest date on that page.
     SEEN     how many times the entry appeared, shown only when duplicates
              were collapsed for the run.

  Each entry is labeled for what it is: Diagnosis, Finding/symptom, Finding,
  Procedure, Event, Situation, Status/encounter, Test order, or Clinical entry.
  An ICD-10 R-code (Snoring, chest pain, dizziness) is something the record
  OBSERVED, not a disease a doctor diagnosed, so it is labeled Finding/symptom.
  It keeps its code and its research result; only the word changes.

  Where a section says the local reference had no result, it says so ONCE at the
  top rather than after every row. In the page-by-page appendix an entry with no
  "Potential VA DC" line is one the reference simply has no row for.
  It is NOT medical or legal advice, NOT an official VA mapping, NOT a rating
  prediction, and NOT a recommendation about what to claim. Confirm every item
  against the source record.

WHAT TO TRUST
  Code + description : reliable. Every code is checked against the official CDC
                       ICD-10-CM list, so stray look-alike tokens are dropped. A
                       bare (no-decimal) code is only kept when its description is
                       written next to it - so device models, IDs and stray tokens
                       do not sneak in.
  CODE "-"           : a diagnosis named in the record with no code at all.
  [ICD-9-CM]         : coded under ICD-9, not ICD-10. Kept for completeness. Do
  [SNOMED CT ...]      not paste these into an ICD-10 lookup tool. With the
                       optional reference files installed, the research section
                       at the end shows what each one maps to in ICD-10-CM. A
                       SNOMED tag also names what the entry IS, such as
                       [SNOMED CT procedure] or [SNOMED CT event].
  Page number        : real for PDF. Word/Text/HTML have no pages, so a line or
                       paragraph number is shown instead; XML shows an entry number.
  Date               : APPROXIMATE - the nearest date to the code. With free-form
                       documents a date cannot be tied to a code with certainty.
                       Always confirm against the source. A date labeled as the
                       patient's date of birth is never used, and a letterhead
                       date repeated in the page banner yields to a date from
                       the body of the page. Old dates are fine: a 1970s visit
                       date is kept, never dropped for its age.

OCR (scanned pages)
  Scanned/image PDF pages are read with OCR automatically; rows found that way are
  marked (OCR). OCR can misread characters, so confirm them against the source.
  A bare category code (no decimal, e.g. L40) from OCR is kept only when its own
  official description is written next to it, so stray marks that happen to spell
  a valid code are still dropped. OCR takes a few seconds per scanned page; a big
  file can take several minutes (progress prints as it goes). Tick "Skip scanned
  pages" to skip it.

NOT DONE
  No old .doc (only .docx).

UPDATING THE CODE LIST
  icd10_data.tsv is the CDC ICD-10-CM code set (code <TAB> description). Replace it
  with a newer year's file in the same format to update.

SCANNING A FOLDER
  Reports this tool has already written are skipped, so re-running on the same
  folder does not read its own output back in as if it were a medical record.

PRIVACY
  Runs entirely on your computer. No internet connection is used. The
  ratemyvso.net address is printed in the report for you to visit yourself - the
  tool never sends anything anywhere.

OPTIONAL REFERENCE FILES BESIDE THE PROGRAM
  dc_reference.json + dc_reference.manifest.json          the Potential-VA-
     Diagnostic-Code research cross-reference. SHIP AS A PAIR or not at all.
  icd9_gem_reference.json + icd9_gem_reference.manifest.json   the ICD-9-CM
     bridge into that same reference. Needs the pair above; SHIP AS A PAIR.
  snomed_icd10cm_reference.json.gz + snomed_icd10cm_reference.manifest.json
     the SNOMED CT map into that same reference, and what each SNOMED code IS.
     SHIP AS A PAIR. About 7 MB, and it is read straight out of the .gz - it is
     never unpacked into a 61 MB file beside the program. It works on its own: if
     the DC pair above is missing you still get the ICD-10-CM translation, and
     only the VA diagnostic code says it is unavailable.
  With any of them missing or damaged the program still extracts everything -
  only the extra research section reports itself unavailable. See the report's
  own wording: it is a research cross-reference, never an official VA mapping.

SNOMED CT CODES (only when the map file above is installed)
  Records often code their problem list in SNOMED CT. Those rows are always
  reported and tagged [SNOMED CT]. With the map installed, the report also shows
  which ICD-10-CM code the licensed SNOMED CT US Edition map gives. The SNOMED
  code in your record's rows never changes, and a mapped code is never added to
  the ICD-10 paste list.

A SNOMED CODE IS NOT ALWAYS A DIAGNOSIS - READ THE TYPE
  SNOMED numbers name diseases, but they also name findings, procedures, events
  and situations. An appendectomy has a SNOMED code. So does "exposure to
  COVID-19", and that one translates cleanly to an ICD-10-CM code - which says
  only that the two coding systems agree on the words. It does NOT say a doctor
  found you had COVID-19.

  So every SNOMED row now says what it is:

     [SNOMED CT disorder]   a disease or condition
     [SNOMED CT finding]    something observed, like chest pain
     [SNOMED CT procedure]  an operation or treatment
     [SNOMED CT event]      something that happened, like an exposure
     [SNOMED CT situation]  a circumstance, like a family history
     [SNOMED CT]            the map could not tell us which

  ONLY A DISORDER is offered a potential VA diagnostic code. A procedure, a
  finding, an event, a situation, or an entry whose type could not be read is
  listed in its own "SNOMED CT clinical concepts" section near the end, with what
  it is and what it translates to, and nothing is suggested for it. Your record's
  entry is never removed or reworded: only the label on it changes.

  If you believe an entry is your diagnosis and the tool calls it something else,
  the tool is repeating what the official SNOMED release says that number means.
  Check the number against your record.

  Read the MAP BASIS column too, because it decides how much the row is worth:

  direct, unconditional
     One clear target. Looked up like any ICD-10 code.
  multiple unconditional map groups
     The map gives SEVERAL ICD-10-CM codes that apply TOGETHER. Each is looked
     up separately and you get more than one answer. That is correct, not a
     contradiction.
  conditional map, review required
     The map cannot finish without something only the record can tell it - a
     patient's age at onset, type 1 versus type 2 diabetes, whether this is a
     first visit or a follow-up. The possible codes and the map's own conditions
     are listed and NOTHING is chosen for you. No diagnostic code is shown,
     on purpose. Decide it against the record yourself.
  no ICD-10-CM classification from this map
     The map says this concept cannot be classified with the data available.
  not in this SNOMED CT to ICD-10-CM map
     The number is not in the map at all.

  A code ending in "?" such as T14.8XX? is not a finished code. The "?" is a
  character the map could not supply (first visit, follow-up, or a lasting
  after-effect). Those rows are always marked for review.

  Only codes the record LABELS "SNOMED" or "SNOMED CT", or an XML entry carrying
  the SNOMED CT code system, are ever crossed. A bare long number in a record is
  usually an account or accession number and is never treated as a diagnosis.

  The map is the licensed SNOMED CT US Edition (March 2026), and so is the list
  of what each code IS. It is a terminology map, not a VA document: the result is
  research, never a rating or a decision.

FILES THIS TOOL WRITES BESIDE ITSELF
  icd_tool_accepted.txt   who accepted the terms, when, and for which version
  icd_tool_settings.txt   the folder you used last
  icd_tool_error.log      only if something went wrong (see below)
  (If this folder is read-only, these go to %LOCALAPPDATA%\ICD_Tool instead.)
  Beside the RECORDS it writes icd10_report.txt and, on request,
  icd10_report_claim_handoff.txt.

IF THE WINDOW MISBEHAVES
  The window version has no console, so errors cannot print anywhere you would
  see them. Instead:

  1. An unexpected error pops up a message box and is appended to
     icd_tool_error.log in this folder. Send that file on.

  2. For a problem with no error at all - nothing appears, or it seems to hang -
     turn on tracing and run it again:

        set ICD_DEBUG=1
        icd_extract_gui.exe

     That writes a timestamped icd_tool_debug.log showing how far startup got.
     Tracing is off unless ICD_DEBUG=1, so it costs nothing the rest of the time.

  Note the first launch after a reboot is slower than later ones: the .exe is a
  single file over a hundred megabytes that unpacks itself to a temporary folder each time it runs.
  Give it a few seconds before assuming it has died.
