Model Assessment
/model-assessmentNEWWhat it does
Validation design, task-correct held-out metrics (Dice + HD95, AUROC + AUPRC, FROC, calibration), uncertainty/OOD and explainability for trained imaging models, each with a gate.
Highlights
- ✓Split-leakage gate proves patient disjointness
- ✓Task-correct metrics with bootstrap CIs, computed only from executed code
- ✓Flags unmeasured uncertainty coverage and an unvalidated OOD guard
- ✓Grad-CAM sanity checks and quantitative localisation vs ground truth
Install this skill
git clone https://github.com/Aperivue/medsci-skills.git
mkdir -p ~/.claude/skills
cp -r medsci-skills/skills/model-assessment ~/.claude/skills/Related skills
Paper-grounded architecture choice for medical imaging, then a check of the actual repo or checkpoint: licence, version pin, weight provenance and benchmark overlap.
Model Scaffold/model-scaffoldGenerate a reproducible, runnable PyTorch training repo for a medical-imaging task — segmentation, classification, detection, synthesis, self-supervised pretraining, or fine-tuning a pretrained backbone — with a patient-level seed-locked split, train/evaluate scripts, and a Methods stub. Integrates MONAI / nnU-Net, never reimplements them.
Model Card & Datasheet/model-cardGenerate the documentation an engineer-built model must carry — a Model Card (Mitchell et al. 2019), a Datasheet for its dataset (Gebru et al. 2021), and a METRIC data-quality pass — filled only from user-supplied facts, then verify every required section is present with a completeness gate.
LLM/MLLM Evaluation/mllm-evalModel-agnostic evaluation harness for an LLM or MLLM on a clinical task — report generation, visual question answering, clinical text extraction — covering the adjudicated reference, clinical-efficacy metrics (RadGraph-F1 / CheXbert-F1), faithfulness, contamination, prompt sensitivity, and a reader study.