Analysis at a glance
A compact transcript panel separated four recorded stimulation conditions in macrophages from donors the model had never seen. The primary compact-panel procedure classified all 28 held-out profiles correctly (balanced accuracy 1.00). The fixed five-gene panel from my thesis scored 0.79, and a training-prior baseline scored 0.25.
The question
My M.D. thesis measured macrophage responses to probiotic bacteria with qPCR and cytokine assays. This is a narrower, separate question asked on public data: how much of a recorded stimulation condition can a model recover for a donor it has never seen?
If a small panel of genes can separate the conditions in donors held out entirely from training, is that panel a candidate worth testing properly? The answer here is a cautious yes, with the limits stated below rather than buried.
The data
The data are the public GSE46903 series from Xue and colleagues. It contains several cell types and stimulation schedules. The question here is narrower: GM-CSF differentiated macrophages sampled 72 hours after one of four recorded conditions.
The series has 384 arrays, including 299 marked macrophage. Requiring one profile of each condition per donor leaves 7 donors and 28 profiles, seven per class. Probes with ambiguous or missing gene symbols are discarded, and probes mapping to a single gene are reduced by a within-sample median, giving 19,235 mapped genes with no missing cells.
These are transcript measurements. TNF and IL1B here are not the secreted proteins my thesis measured by ELISA.
Holding donors out
Samples from the same donor share biology and laboratory history, so they stay together in one fold. Feature ranking, imputation, scaling and model tuning use training donors only, and an outer held-out donor is predicted once. There are three outer folds and two inner folds for tuning.
With seven donors the training groups are small, four or five donors each. That is the weakest part of the design, and no amount of care in the pipeline removes it.
This matters more than it first appears. A random split of the 28 profiles would score higher, but it would also put the same person's cells on both sides of the split, which flatters the result and means nothing for a new donor.
Which genes are enough?
The fixed panel is ABCA1, ABCG1 and NR1H3, the
gene for LXR-alpha, alongside TNF and IL1B. The fixed panel comes from
measurements made in a previous thesis, so it is a fair benchmark rather than a tuned one.
A separate compact-panel procedure ranks genes inside each training fold and chooses among 5, 10, 20 and 50 genes using the inner folds. These are computational panels, not validated assays, and the selected panel differs between folds.
Training a small network in PyTorch
The same ten-gene input feeds a network of 10 → 8 → 4 units: an
eight-unit ReLU hidden layer and four output scores, one per condition. That is
124 parameters in total. The whole thing trains on one CPU thread in
about forty seconds, and the model itself occupies under two kilobytes of memory —
the library is the expensive part, not the network.
It is trained with full-batch AdamW for a fixed 150 epochs under cross-entropy, and the schedule is set in code before any evaluation rather than chosen afterwards. No parameter was tuned against held-out results, so the score below is a single pre-registered attempt rather than the best of several.
Training loss per fold
| Epoch | Fold 0 | Fold 1 | Fold 2 |
|---|---|---|---|
| 1 | 1.5445 | 1.3555 | 1.5285 |
| 10 | 1.1583 | 1.0435 | 1.0881 |
| 25 | 0.4860 | 0.4079 | 0.4765 |
| 50 | 0.1105 | 0.0776 | 0.2218 |
| 100 | 0.0077 | 0.0114 | 0.0116 |
| 150 | 0.0042 | 0.0061 | 0.0052 |
The loss falls to roughly 0.005 by epoch 150. On twenty-one training profiles that is close to memorisation, which is the honest reading: the schedule has no early stopping and no held-out monitoring. It happened not to hurt the held-out score, and the fact that regularized logistic regression reaches the same 1.00 on the same genes is the reason to read this as an easy task rather than a successful architecture.
This is a coding and comparison experiment on real biological data. It shows that a network can be built, trained and evaluated reproducibly under a proper held-out protocol. It does not show that deep learning suits this task, and it says nothing about imaging, longitudinal records or clinical validation.
A held-out sample
How the models did
| Procedure | Balanced accuracy | Donor bootstrap interval |
|---|---|---|
| dummy | 0.25 | 0.25–0.25 |
| thesis lr | 0.79 | 0.57–0.96 |
| compact lr | 1.00 | 1.00–1.00 |
| compact lr k5 | 0.93 | 0.82–1.00 |
| compact lr k10 | 1.00 | 1.00–1.00 |
| compact lr k20 | 1.00 | 1.00–1.00 |
| compact lr k50 | 1.00 | 1.00–1.00 |
| broad lr | 1.00 | 1.00–1.00 |
| compact gb | 0.93 | 0.86–1.00 |
| torch mlp k10 | 1.00 | 1.00–1.00 |
These intervals resample stored donor predictions and do not repeat training. A 1.00–1.00 interval means this tiny set contained no errors; it does not imply that a new cohort would have near-perfect accuracy.
Primary compact-panel confusion matrix
| Recorded \ Predicted | con | IFNg | IL4 | TNFa+PGE2+P3C |
|---|---|---|---|---|
| con | 7 | 0 | 0 | 0 |
| IFNg | 0 | 7 | 0 | 0 |
| IL4 | 0 | 0 | 7 | 0 |
| TNFa+PGE2+P3C | 0 | 0 | 0 | 7 |
Does the model know when it is wrong
A model that classifies every held-out sample correctly has told you nothing about whether it would know when it was wrong. The network scores 1.00 here, so its own calibration cannot be measured: with no errors the likelihood has no interior optimum, and a fitted temperature runs to its search boundary. The measurement was moved to the five-gene panel, which does make errors, using its already-saved predictions.
| Quantity | Value | What it means |
|---|---|---|
| Mean confidence when correct | 0.636 | How sure it is when it is right |
| Mean confidence when wrong | 0.393 | How sure it is when it is wrong |
| Margin between top two classes, correct | 0.416 | A clear separation |
| Margin between top two classes, wrong | 0.090 | The model is barely deciding |
| Expected calibration error | 0.201 | Gap between stated confidence and actual accuracy |
The margin is the difference between the two highest class probabilities. It falls sharply on the samples the panel gets wrong, which is the only reason to show a confidence value to anyone: a number that does not track correctness is decoration.
The errors are not spread evenly. They fall in 3 of 7 donors, and donor BC58 alone accounts for 3 of them. A per-sample confidence would mark the affected samples as uncertain, but it would not say which donor is unrepresented in a five-gene panel. That is a limitation of the panel, not of the confidence estimate.
Computed from the frozen prediction file, not refitted. The expected calibration error says the panel is over-confident overall; the margin gap is what says the confidence carries per-sample information.
What this can tell us
These samples come from one experimental study and seven matched donors. The matrix was already quantile-normalised when deposited, potentially across the arrays later assigned to different folds. That upstream step cannot be undone by a training-fold pipeline, so this describes classification within an already-processed research dataset.
The conditions are known laboratory exposures. A prediction here is not a macrophage diagnosis, a measure of cytokine secretion, or a treatment recommendation. A selected gene may be useful to the model without being a causal regulator.
A 1.00 bootstrap interval over a perfect set of stored predictions says nothing precise about accuracy in a new population. Testing this panel as an assay would need new donors and separate experiments.
Reproduce and inspect
The series matrix and platform annotation are public, and the code records input checksums and the chosen sample IDs. Download the run record, donor splits, selected genes, task definition, the PyTorch training trace and uncertainty report so a reader can check the outcome rather than trust a screenshot.
The training code is served as source: torch_experiment.py and uncertainty.py.
That project uses single-cell data from carotid plaque. This one uses a different assay and a different cohort. Neither trains on the other.