A nose for cancer

Brush the inside of a nose, image the cells without a single stain, and their shapes alone carry enough to separate cancer from inflammation. A new study suggests the nose has been telling us this all along.

A sinonasal cancer and a bad sinus infection can look like the same illness. Both block the nose and sinuses, and can be difficult to easily differentiate. The consequences of that resemblance are not small. Two-thirds of sinonasal squamous-cell carcinomas are found only once they have reached an advanced stage, by which point the surgery is more complicated, the odds are worse, and the cancer has had months to develop hiding behind symptoms no one had reason to question and therefore biopsy.

The tools that could resolve the ambiguity earlier are the ones patients are least eager to receive: a scope, a biopsy, an MRI. So most of the time no one reaches for them, and the cancer keeps its cover. The nose itself is easy to sample — a soft brush against the inside of a nostril collects thousands of cells in seconds, the same trick that made the Pap smear one of the most successful screening tools in medicine. The problem was never getting a cell sample. It was reading them once you had.

A new study in npj Precision Oncology, led by Brittany Rupp and Kevin Byrd at Virginia Commonwealth University with collaborators at the University of North Carolina, takes a serious run at that reading problem. Using the REM-I platform, the team imaged more than 2.5 million single cells from nasal brushings across three groups of patients: healthy volunteers, people with chronic rhinosinusitis, and people with biopsy-confirmed sinonasal carcinoma. No stains. No antibodies. No genetic sequencing. Just high-speed brightfield imaging of unlabeled cells, and an AI model was asked to find the differences between a nose that is inflamed and a one that is turning malignant.

It found one.

The pipeline: a nasal brush, label-free imaging on REM-I, morpholomic features read by AI, then unsupervised clustering. Adapted from Rupp et al., npj Precision Oncology (2026). CC BY-NC-ND 4.0.

Comte’s mistake

In 1835 the philosopher Auguste Comte reached for an example of something humanity could never possibly know. He chose the chemical composition of the stars. They were too far away; we would only ever have their light, and light, he assumed, could not tell you what a thing was made of. Within a few decades he was comprehensively wrong. Fraunhofer had already noticed dark lines ruled across the solar spectrum without knowing what they meant; Kirchhoff and Bunsen worked out that each line was a fingerprint left by a particular element. The composition of the stars turned out to be written into exactly the signal Comte thought was empty. We had been staring at it the whole time. We simply could not read it.

A cell under a microscope is a similar kind of signal. Its shape, its texture, the way light scatters through its interior — all of that is a low-resolution projection of an enormous hidden state: which genes are switched on, what proteins have accumulated, how the cytoskeleton is arranged, whether the machinery has begun to misbehave. Cytology has read that projection by eye for a hundred years, ever since Papanicolaou realized that abnormal cervical cells announce themselves in their appearance before anything else. But the human eye reads only the coarse lines of the spectrum: size, roundness, a suggestion of granularity. Most of the information in the picture has stayed unread, not because it was absent, but because we lacked the tools to decode it.

Immune cells imaged from a healthy nose, a tumor, and chronic inflammation. To the eye the three are hard to tell apart; the signal that separates them lives in features no pathologist names. Adapted from Rupp et al., npj Precision Oncology (2026). CC BY-NC-ND 4.0.

But now we have a model. For each of the 2.5 million cells, REM-I extracts 115 morphological features — 51 that a person could name and measure, and 64 learned directly from the raw images by our Human Foundation Model, features that correspond to no textbook description of a cell. When the team asked which features best separated cancer from inflammation, the answer was consistent and slightly humbling: the ones not interpretable by a human! The signal lives largely in features no microscope makes legible to the eye — not because pathologists miss anything, but because a century of cytology has had to read morphology one slide at a time. The model reads the same kind of information, quantified across millions of cells. 

Reading without a dictionary

There is a catch worth being honest about, this catch shadows every deep-learning result in biology. Knowing that a feature(s) separates cancer from inflammation is not the same as knowing what it means. The model can classify the spectrum before it can translate it.

The team did not wave this away. They correlated the learned features against the interpretable ones and found that many of them track composite, biologically sensible programs — cell size bundled with texture, boundary irregularity, the unevenness of internal intensity. One feature the human eye can actually grasp did much of the work in the immune compartment: small dark pixel intensity, essentially a count of tiny dense specks inside a cell. It rises with the granules that myeloid cells carry, which makes it a decent proxy for how active the innate immune response has become. It is a small thing, a measure of specks, and it turns out to carry a real part of the diagnostic load.

To give the learned features something to stand on, the team first built a reference atlas the field has not had: more than 641,000 images of purified immune cells, seven types spanning the myeloid and lymphoid lineages, each imaged under the same conditions so the model could learn what a known cell looks like before being asked to name an unknown one. Immune lineage, it turns out, is strongly written into morphology. Myeloid cells are more distinguishable by shape than lymphoid ones, which echoes exactly what a cytologist with a stain has always known, arrived at here without the stain.

The company a tumor keeps

Perhaps the most interesting result is not about the cancer cells, but about their neighbors.

When the team looked at the immune cells scattered through a tumor swab, they found the tumor’s signature written there too. Malignant samples carried more myeloid cells than healthy ones, though not reliably more than inflamed ones. What set them apart was activity rather than headcount: the granule signal inside those myeloid cells ran higher in cancer than in either a healthy nose or an inflamed one. A population of basophil- and NK-enriched cells picked up learned morphological features that appeared in no healthy or inflamed sample — a shape that, in this cohort, showed up only alongside the cancer. The tumor epithelium, meanwhile, told its own story: the malignant cells were consistently smaller, with a distinct textural fingerprint, matching what histopathology has long described in stained tumor tissue.

This is the quietly important part. The disease shows up across the entire cellular milieu that a swab captures, immune and epithelial cells alike. A tumor changes the neighborhood it grows in, and a label-free image can see the neighborhood change. Combine the immune and epithelial signals and you have something closer to a portrait of the tissue’s state than any single marker would give you.

Signal, then scale

The team went one step further and built a predictor from four features, two immune and two epithelial. Tested by leave-one-out cross-validation, it separated tumor from non-tumor swabs with an area under the curve of 0.94, catching every cancer and calling one inflamed patient a tumor.

Take the 0.94 with the salt the authors themselves offer. Those four features were hand-picked from tests run on the same handful of patients the model was then scored against, which is a reliable way to flatter a number. It says something is there. It does not say how well anything works. The paper concedes the point and names the remedy: more patients, and a feature selection that does not get to peek.

A four-feature model separates tumor from non-tumor swabs with an AUC of 0.94, catching every cancer in a small first cohort. Adapted from Rupp et al., npj Precision Oncology (2026). CC BY-NC-ND 4.0.

Three cancers, five inflamed, five healthy, and every tumor the same subtype of squamous-cell carcinoma. The cancer patients were also about a quarter-century older than everyone else, and all of them were men, a skew the authors put on the page themselves rather than bury. Read the result as an invitation to a larger study, not a finished test. That is the ordinary path of a promising first result: establish that the signal is there, then go and assemble the cohort that turns it into a diagnostic.

What the study already establishes is the durable part: that the information needed to tell these two conditions apart survives in unstained single-cell morphology, and that a model can recover it. That is worth taking seriously now. The patients who will confirm it are the next chapter.

What comes after the swab

Everything above is what the study shows. What follows is where we think it leads, which is a different sort of claim and worth labelling as one. The instrument and the model are ours; the patients, the study, and the clinical judgment belong to the academic team. What makes the result worth considering is how general it is. The same label-free approach has shown early promise in melanoma and in cancer cells that have learned to resist a drug, pulling states out of morphology that were supposed to need a molecular assay. Nasal cytology is one more instance of a pattern that keeps recurring: high-dimensional shape is a richer readout of cell state than we have been treating it as.

The natural next move is to stop reading and start sorting. Because REM-I identifies these cells in real time, a morphologically defined population — the tumor-associated immune cells, say — can in principle be pulled aside while still intact and handed to sequencing, which would begin to translate the learned features back into biology. Classification, then decipherment. The order matters, and it is the order the history of spectroscopy already ran once.

Comte was wrong about the stars because he mistook the limits of his instruments for the limits of nature. The information he wanted was in the light all along. A cell’s image is the same kind of light. It has been carrying the state of the cell in every picture we have ever taken, waiting on an instrument patient enough to read the lines. Cellular morphology has always told a story and now we have the means to discover it. 

Next
Next

Shape of My Cell: Reading the morphological "drift" of the cellular factory