Student research · Explainable AI for medical imaging
Transparent AI for Medical Reasoning.
Founded in 2024 by two student researchers, Aletheia pairs deep-learning classifiers for blood smears and brain MRI with Grad-CAM heatmaps. A clinician can see which parts of a scan drove a prediction instead of trusting a black box.
Each pipeline classifies a scan and then explains its answer with a heatmap over the original image.
Blood smears
Leukemia staging
Sorts a blood-smear image into four classes: Benign, or one of three stages of acute lymphoblastic leukemia (Early, Pre, Pro). The network is MobileNetV2, pre-trained on ImageNet and fine-tuned on our smear images.
Input: 224 × 224 RGB images
Training: Adam (learning rate 1e-4), class-balanced weights, flip, rotation, zoom, brightness and contrast augmentation
Explanation: Grad-CAM on the final convolutional layer
Best validation accuracy: — (preliminary)Training log
Brain MRI
Alzheimer's staging
Sorts a 2D brain MRI slice into four stages: non-demented, very mild, mild and moderate. A compact convolutional network does the classifying. The study's real question is interpretability: does the model look at the anatomy that changes with the disease?
Data: Hugging Face Falah/Alzheimer_MRI, 128 × 128 grayscale slices
Split: 80/20 stratified train and validation
Model: two convolutional layers (32 and 64 filters), max pooling, 30% dropout
The network reads the scan and outputs a probability for each class. The top class is its answer.
2. Trace the decision
Grad-CAM measures how strongly each region of the network's last convolutional layer pushed toward that answer.
3. Show it
Those influence scores become a heatmap laid over the scan. Warm colors mark the regions that mattered most.
Built on consumer hardware
We trained in Google Colab and on a local desktop with an Intel Arc A750 GPU. Both networks are small by modern standards. Explainable AI shouldn't need a data center.
Behind every stained biopsy slide and blurred neuroimaging scan is a human life hanging in the balance. We built Aletheia because waiting in the shadow of diagnostic uncertainty is an unacceptable failure of modern medicine.
The human cost of delay
The cruel mathematics of late detection
In high-stakes medicine, whether tracking aggressive acute lymphoblastic leukemia or charting the silent structural creep of Alzheimer's, time is quite literally tissue, memory, and life.
When a diagnosis is delayed, miscalculated, or hidden behind the impenetrable black box of a traditional algorithm, the consequences ripple outward. Late detection robs patients of therapeutic windows and turns treatable conditions into irreversible ones. Families are left wondering whether an earlier look beneath the cellular surface could have changed the outcome.
Medicine does not fail because doctors lack compassion. It fails when diagnostic tools lack the clarity, speed, and transparency needed to catch micro-malignancies and early-stage degeneration before they entrench themselves.
Accountability and explainable intelligence
Bridging the chasm between AI and clinical trust
As artificial intelligence moves into the clinic, software increasingly outputs high-stakes predictions without explaining why. Algorithms cannot bear moral responsibility, so human experts must keep rigorous oversight of every automated decision.
Yet how can a physician stake a patient's life on an algorithm they cannot audit? If an AI tells a pathologist a slide is malignant without pointing to the chromatin abnormalities or nuclear distortions behind that call, it is a digital oracle, and it is unfit for the gravity of the ward.
Visible reasoning
Grad-CAM heatmaps show which regions of a scan triggered a classification, so the reasoning can be checked.
Reproducible evidence
We publish our training logs, confusion matrix and limitations so others can scrutinize the results.
Physician in charge
The tool supports a clinician's judgment. It never replaces it.
Our promise to every patient and clinician
We want no disease to hide in the blind spots of technology. Transparency isn't just an engineering choice. It is a moral obligation.
Build a made-up scan, move structures around, then switch on the attention map. The heatmap follows whatever you place, the same way Grad-CAM follows what a network finds important. The results here come from simple rules, not from our trained models, and this is not a diagnostic tool.
Scan
Add structures
Drag one onto the scan, or select it to add it near the center.
Workspace · click a structure to select it, drag to move it
Structures
Attention scale
Low influenceHigh influence
Benign structures glow faintly. Abnormal ones dominate the map.
Simulated reading of this scene
Simulation, not a model output
Pattern in this scene—
Attention on abnormal structures—
Strongest attention—
How the real heatmaps are made and how well the real models perform: Credibility & Research.
Creators and references
Credits & Citations.
Project creators
PY
Paul Yoon
Lead programmer · age 17
Paul is a hardworking student athlete with an intense passion for medicine and artificial intelligence. As lead programmer he writes the project's code, from model training to this website.
KD
Kyle Dirius
Lead researcher & analyst · age 17
Kyle is a dedicated student athlete driven by a lifelong curiosity for clinical medicine and machine learning. He leads experimental design, evaluation, and dataset curation so that every result rests on sound evidence.
References
Alzheimer's data: Falah, Falah/Alzheimer_MRI, Hugging Face dataset hub.
Leukemia data: four-class blood-smear image set (Benign, Early, Pre, Pro).
Grad-CAM: Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. IEEE International Conference on Computer Vision (ICCV).
MobileNetV2: Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L.-C. (2018). MobileNetV2: Inverted residuals and linear bottlenecks. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
Tools: TensorFlow / Keras, Google Colab, Tailwind CSS, Font Awesome.
Hardware: local desktop with an Intel Arc A750 GPU, plus Google Colab.
Empirical research
Credibility & Research.
What we trained, how it performed, and where the results stop being trustworthy. Numbers on this page come from our own training runs.
Study design
Staging acute lymphoblastic leukemia with a pre-trained network
The question: can an ImageNet-pre-trained network separate benign blood cells from three stages of acute lymphoblastic leukemia (Early, Pre, Pro), and can Grad-CAM show what it relied on?
The model is MobileNetV2 with its ImageNet weights frozen, followed by global average pooling, a 512-unit dense layer and a four-way softmax. Images are resized to 224 × 224 RGB. We trained with Adam (learning rate 1e-4), categorical cross-entropy and class-balanced weights, using random flip, rotation, zoom, brightness and contrast augmentation. Training ran for five epochs, keeping the checkpoint with the best validation accuracy and using early stopping and learning-rate reduction on plateau. The split was 80/20 train and validation, with about 2,600 training images (82 batches of 32).
Grad-CAM is computed on the final convolutional layer (Conv_1, a 7 × 7 grid of 1,280 feature maps), upsampled and overlaid on the input image.
Training log
Accuracy by epoch
Validation accuracy peaked at — in epoch —, and the weights from that epoch were kept.
Reading this number carefully: the same validation set picked the best epoch and produced the headline accuracy, so — is an optimistic, preliminary estimate. A separate held-out test set and a full confusion matrix are the next step.
Study design
Interpretable Alzheimer's staging from MRI
The question: to what extent can a convolutional neural network paired with Grad-CAM heatmaps make Alzheimer's diagnosis from MRI scans more interpretable?
We used the Hugging Face Falah/Alzheimer_MRI dataset: 6,400 grayscale slices at 128 × 128 pixels across four stages, non-demented (noimp), very mild (verymildimp), mild (mildimp) and moderate (modimp). The data were split 80/20, stratified by stage.
Overfitting is the main risk, because a model can memorize imaging artifacts instead of disease markers. The network therefore stays small: two 3 × 3 convolutional layers (32 and 64 filters), max pooling, and a 30% dropout layer before the classifier.
Quantitative validation
Confusion matrix
Rows are the true stage and columns the predicted stage. The diagonal shows correct predictions.
Visual interpretability
Grad-CAM by stage
Warm colors (red, yellow) mark regions with high influence on the prediction. Cool colors (blue) mark regions that contributed little.
Non-demented
Activation is spread across outer cortical regions with no concentrated hot spot, as expected for a healthy baseline.
Very mild
Activation concentrates near the center of the brain, around subcortical structures where early change is expected.
Mild
Activation becomes more distinct across cortical and subcortical regions, where hippocampal atrophy and ventricular enlargement develop.
Moderate
Pronounced central and temporal activation, consistent with advanced structural degeneration.
Limitations and responsible use
Aletheia is a student research prototype. It is not a medical device and must not be used to diagnose or treat anyone.
Both models were trained on public datasets. How they behave on other scanners, staining methods or patient populations is unknown.
A heatmap shows where a network looked, not whether it was right. A highlighted region is not proof of disease.
Reported metrics are preliminary. They come from a single train/validation split, not an independent test set.
The live engine uses synthetic scenes and simple rules. It illustrates the idea of an attention map and does not run our trained models.