Molecular brain imaging during gender‑affirming hormone therapy: a pilot synthesis of PET and MR spectroscopy
Counted once
What has been measured, in whom, and what a focused paper could add.
PET · MR spectroscopy · selected MRI / five PET reports, three cohort families, two longitudinal MRS cohorts
The short answer
A focused synthesis of PET and MR spectroscopy during gender-affirming hormone therapy (GAHT) looks more promising than a PET-only review. The treatment PET evidence comes from only two Vienna cohort families and two neurochemical targets, so it cannot support a target-level meta-analysis. MRS adds different measurements and one clearly separate longitudinal cohort, from Ghent.
The contribution such a paper could make lies less in collecting studies than in counting them properly. That means separating reports from cohorts and visits from people, and separating follow-up observations from people with a valid measurement both before and after treatment. Then each molecular signal can be read for what it does and does not measure. That is the thread this report follows, and the reason for its title.
5
reportsPET publications identified, including one perfusion study
3
familiescohort families behind those five reports
2
familiesbehind the three PET reports that analyse change during treatment
11
pairedtransgender participants with both harmine PET scans (Handschuh 2024)
How to read this reportEvery number carries its unit
The commonest error in this literature is a unit error: adding visits as if they were people, or papers as if they were cohorts. So every load-bearing number here is underlined in the colour of what it counts. Hover over or tap any of them.
□ reports◇ cohort families● people●● paired people (before and after)○ observations / scans∙ search records
Small chips mark what kind of claim a sentence makes: an result, the interpretation, reasoning across sources, a , or an question that needs more source work. The dots show how deeply the source behind the claim was read. Citations such as Kranz 2015 open a summary card. Defined terms such as binding potential carry a soft highlight. Brief mode, in the top bar, keeps only headings, key paragraphs, figures and tables.
Five distinctionsWhat this report keeps apart
Five separations organise everything that follows. Each one prevents a specific, common overstatement.
A paper is not a cohort. Three PET papers analysing treatment come from two recruitment families, and five MRS papers come from two longitudinal families plus one cross-sectional cohort.
A visit is not a person. The DASB study measured 25 transgender participants at four weeks and 24 at four months, and these are largely the same people.
A follow-up observation is not a paired change.Ten usable four-month harmine scans in Kranz 2021 are not a verified ten before-and-after pairs.
Binding is not serotonin, and a ratio is not a concentration. Binding potential, distribution volume and metabolite ratios measure different things, and none of them alone establishes clinical benefit, harm or anything about gender identity.
A second analysis of the same people is not a replication. The 2021 and 2024 harmine papers ask different regional questions of one study family, so their different results neither confirm nor refute each other.
The four sectionsA short route through the evidence
Each section opens with the same four lines: what was measured, in whom, what it can mean, and what would settle it. The sections then earn those lines.
Two molecular targets, each resting on one Vienna family
Treatment PET during gender-affirming hormone therapy has measured two proteins, the serotonin transporter and monoamine oxidase A, each in a single Vienna study family of a few dozen people. Transporter binding rose under testosterone. The monoamine oxidase results run from regional decreases to no significant change, but they are two analyses of one cohort, not a replication.
Real structure: human MAO-A with harmine, PDB 2Z5X · Photon pairs schematic
MeasuredHow much of a radiotracer two proteins hold: SERT with [11C]DASB (binding potential) and MAO-A with [11C]harmine (distribution volume).
In whomOne Vienna family per target: 33 transgender participants at the SERT baseline, and 11 with both MAO-A scans in the latest report.
What it can meanRegional target availability changed during treatment in small samples. Nothing direct about synaptic serotonin, mood or gender identity.
What would settle itAn independent cohort per target, reported paired counts, regions fixed in advance, and the planned [11C]AMT data, if they exist.
MeasurementWhat a PET number means here
Both targets belong to the serotonin system, and both measurements count tracer held in tissue. They differ in what that count is compared with, and that sets how far each number can be read.
A positron emission tomography (PET) study starts with a radiotracer: a drug-like molecule that carries a short-lived positron emitter and is injected in trace amounts. Here the emitter is carbon-11, with a half-life of 20.4 minutes, so each dose is made on site shortly before the scan. [11C]DASB binds the serotonin transporter (SERT), the membrane protein that clears released serotonin and that SSRI antidepressants block. [11C]harmine binds monoamine oxidase A (MAO-A), the mitochondrial enzyme that breaks serotonin and other monoamines down. In both Vienna programmes each scan ran for 90 minutes on the same full-ring scanner.
Figure 1.1The two targets, at the same scale. Both proteins are drawn from crystal coordinates and cut open through the bound molecule: SERT with the antidepressant paroxetine in its central site (Coleman 2016), and MAO-A with harmine beside its flavin cofactor (Son 2008). Membrane positions come from the OPM database (Lomize 2012). Each atom is a sphere, and pale areas are the cut faces.
RCSB PDB 5I6X and 2Z5X, membrane frames from OPM; rendered in code for this report.No structure with [11C]DASB bound was used. A crystal structure shows where a molecule sits in purified protein, not how much protein a living brain contains.
Binding potential (BPND) · SERT
Specific [11C]DASB binding relative to a reference region, cerebellar grey matter, which is assumed to hold almost no SERT. No arterial blood is drawn. Can mean: more or fewer transporters available to the tracer. Cannot mean: serotonin concentration. It cannot separate density from affinity, and it moves if the reference region moves.
Distribution volume (VT) · MAO-A
Total [11C]harmine in tissue relative to intact tracer in arterial plasma, measured through an arterial line with metabolite correction. It includes nonspecific binding. Can mean: more or less enzyme available to the tracer. Cannot mean: enzyme activity, serotonin turnover or clinical response.
Neither number is a reading of serotonin in the synapse. The authors of the SERT study suggest that testosterone raised serotonin, which then raised transporter expression . That is a mechanistic proposal, and the scan does not test it.
InventoryFive reports, three families, two targets
Five PET reports include transgender participants. Four measure a molecular target, three analyse change during treatment, and behind those three stand only two recruitment families.
4
reportsmeasure a neurochemical target (SERT or MAO-A)
2
familiesstand behind them, both in Vienna, one per target
24
peopletransgender participants with a four-month SERT scan
11
pairedtransgender participants with both MAO-A scans (Handschuh 2024)
Target by report. Counts are transgender participants or scans exactly as each report gives them; reports in the same family share people and must not be added. Full cohort accounting for PET and MRS together is in section 03.
Report
Target · tracer
Number reported
Design
Transgender participants
Family
Berglund 2008
No molecular target: blood flow · H2[15O]
Activation (perfusion)
Before treatment; odour task
12 TW
Stockholm
Kranz 2014
SERT · [11C]DASB
BPND asymmetry
Before treatment
14 TW, reused in 2015
Vienna SERT
Kranz 2015
SERT · [11C]DASB
BPND
Baseline, 4 weeks, 4 months
33 → 25 → 24
Vienna SERT
Kranz 2021
MAO-A · [11C]harmine
VT, 12 atlas regions
Baseline, about 4 months
18 enrolled; 16 → 10 usable scans
Vienna MAO-A
Handschuh 2024
MAO-A · [11C]harmine
VT, MRI-defined regions
Baseline, about 4.5 months
18 at baseline; 11 paired
Vienna MAO-A
Registry record only
Serotonin synthesis · [11C]AMT
Planned
Planned beside harmine
No verified scans or results
Outside the tally
Searched, not found
Steroid receptors, SV2A, dopamine, TSPO, FDG
None
None
None verified in this pilot
Not proof of absence
Broader reviews place these reports beside structural and functional MRI (Kranz 2020; Nielsen 2025). This section narrows the lens to the molecular targets and what their numbers can carry. The full count of people across PET and spectroscopy, including the open questions about overlap, is in section 03.
SERT · [11C]DASBTransporter binding rose under testosterone
In the only longitudinal SERT study, binding rose over four months of testosterone in transgender men (TM). It fell in a few cortical regions after four months of estradiol plus antiandrogen in transgender women (TW).
The story begins with Kranz 2014, a baseline study of left-right asymmetry in SERT binding in 14 TW and cisgender controls before any treatment. Its methods section later states plainly that this group is a subsample of the 2015 study, so the two papers form one family and their participants cannot be added.
Kranz 2015 scanned 14 TM and 19 TW before treatment, with 35 cisgender controls. Subsets returned after four weeks (9 TM + 16 TW) and four months (11 TM + 13 TW). TM received 1,000 mg testosterone undecanoate every 12 weeks. TW received cyproterone acetate or triptorelin with oral or transdermal estradiol, and six also took finasteride.
BaselinePET 1
4 weeksPET 2
4 monthsPET 3
TMtestosterone
TWestradiol + antiandrogen
ControlsCW + CM, untreated
8 CM rescanned after about 10 days
no 4-month control scans
TotalTM + TW
33 people
25 people
24 people
Figure 1.2Three visits, overlapping people. Each dot is one scan. Solid dots are baseline scans, and every follow-up scan (open dot) belongs to someone already scanned at baseline. The paper does not report how the four-week and four-month subsets overlap.
Kranz 2015, Methods (study design) and Table 1.Dots are arranged by count only; they do not track individual people, whose visit histories are not published.
In TM, binding rose in the amygdala, caudate, putamen and median raphe nucleus, and only the amygdala and caudate survived correction for the number of tests. Larger rises in testosterone went with larger rises in binding at four weeks, but not at four months. In TW, binding fell at four months in the insula and in the anterior and mid-cingulate cortex, and only the anterior cingulate survived correction. Regions lost less binding where estradiol rose more, which the authors read as a protective effect.
The control comparison is not time-matchedOnly 8 cisgender men were scanned twice, about ten days apart (10.13 ± 9.66 days), and their binding was stable. That shows test-retest reliability. It does not show what happens to binding over four months without treatment. Twelve of the 33 transgender participants also had a past mood or anxiety diagnosis, and menstrual phase was not controlled.
MAO-A · [11C]harmineOne cohort, two analyses, two answers
The MAO-A evidence comes from one registered Vienna trial (NCT02715232). Its first report found regional decreases in TM. A later report on the same PET subgroup asked about different regions and found no significant change.
Kranz 2021 calls itself a preliminary report. The trial enrolled 11 TM, 7 TW and 17 controls before the scanner was irreparably damaged in December 2019, leaving 52 usable scans. Its results table counts 10 TM and 6 TW at baseline and 6 TM and 4 TW at about four months. Distribution volume was computed with an arterial input function in 12 regions chosen in advance from the SERT study. In TM, six regions fell by 8 to 10.5%. In TW, no region changed significantly.
White bars are the baseline scans of 10 TM.
Black bars are the four-month scans of 6 TM. The paper does not say how many of these six are among the ten.
Six starred regions passed post hoc tests inside a global model with p = .043. With depression scores as a covariate the global effect gave p = .079, although the regional results held.
Bars and whiskers are means ± SD of two groups of different size, so the chart cannot show change within a person.
Figure 1.3The MAO-A decrease, as published. The original panel is reproduced with its pixels unchanged; the numbered markers are added in code and explained alongside.
Kranz 2021, Fig. 2c, CC BY 4.0. Counts from its Table 2.Asterisks mark p < 0.05 for individual regions; the paired subset and effect sizes for change within a person are not reported.
Two further details shape the reading. The controls' distribution volume was on average higher at their second scan, so the comparison group itself moved. And the rise in testosterone did not correlate with the change in MAO-A in any region after correction.
Handschuh 2024 returns to the same trial: the same registration, ethics number, PET subgroup sizes and scanner failure. This time the PET regions were those where grey-matter density and diffusivity had both changed during treatment, among them the fusiform and lingual gyri, insula, cingulum and cerebellum. No main or interaction effect on MAO-A distribution volume was significant. Seventeen people completed both PET scans, 11 of them transgender (6 TM and 5 TW).
Kranz 2021
Twelve atlas regions chosen in advance. Usable scans: 33 baseline + 19 follow-up. Result: decreases of 8 to 10.5% in six regions in TM; none in TW.
Handschuh 2024
Regions defined by MRI change in the same people. Usable scans: 35 baseline + 17 follow-up. Result: no significant MAO-A change in any group.
Not a failed replicationThe two reports share NCT02715232 and largely the same people, and they test different regions with different models. The later null neither confirms nor refutes the earlier decreases. Both reports count 52 scans, but they divide them differently between baseline and follow-up, and only the authors can reconcile that.
Outside the tallyWhat PET has not measured, or not yet reported
One more PET study measured blood flow rather than a molecule. One planned tracer has no verified output, and several plausible targets did not turn up at all.
Berglund 2008 scanned 12 untreated TW and 24 cisgender controls with H2[15O] water while they smelled odorous steroids. Its twelve scans per person are task conditions within one session, not treatment visits, and 22 of its 24 controls came from earlier odour studies. It is a separate Stockholm programme and counts only in the all-PET total.
The MAO-A trial registration and its grant also planned [11C]AMT, a tracer for serotonin synthesis, at baseline and four months. Registry records, the grant's final report and its publication list give no transgender AMT scan count and no result. The scanner failure explains why the harmine dataset stops, but it does not establish what happened to the AMT arm. It stays an open question, outside the measured-target tally.
Finally, this pilot's searches found no transgender brain PET with direct steroid-receptor tracers, the synaptic-density marker SV2A, dopamine ligands, TSPO for neuroinflammation, or FDG. The searches were bounded, and a database hit count is not an eligibility count, so this is a search finding, not proof that no such study exists.
AttributionWhy a change cannot yet be pinned on one hormone
Even where binding changed, five features of these designs stop the change from being assigned cleanly to testosterone or to estradiol.
Treatment packages. TW received an antiandrogen, cyproterone acetate (which is also a progestogen) or triptorelin, alongside oral or transdermal estradiol, and some also took finasteride. The authors propose that testosterone acts partly through conversion to estradiol.
Attrition. Follow-up subsets were smaller than baseline, the four-week and four-month subsets differ, and a scanner failure cut the harmine study short.
Modelling. DASB relies on a reference region, while harmine relies on arterial input. Each brings its own sources of error, and regions came from an atlas in one analysis and from MRI change in another.
Multiplicity. Twelve regions, two or three groups and up to three visits produce many tests, and the reports handle correction differently.
Small samples. Each arm had between 4 and 16 people at follow-up, and results are group means of groups with different members.
What would turn signals into estimatesA second, independent cohort for each target, with regions fixed in advance, the paired count stated, controls rescanned over the same interval, and individual regimens reported. A synthesis can ask for exactly these items; section 04 returns to them.
What this section does not show
That hormone therapy raises or lowers serotonin in the synapse; neither tracer measures it.
Any clinical benefit, harm or effect on mood.
Anything about gender identity or its causes.
A replicated or pooled effect for either target.
How many unique people the two Vienna programmes include together; no participant crosswalk was found.
Whether [11C]AMT scans of transgender participants exist.
Two longitudinal spectroscopy cohorts, and only one of them is new
Spectroscopy has followed transgender men through testosterone in two places. Vienna's data come from the same trial as the newer PET; Ghent is a separately recruited cohort. Both report ratios of tissue metabolite pools, not neurotransmitter release.
Code-rendered spectrum · peak positions from the literature (ppm) · Schematic shapes and motion
MeasuredMetabolite signals from a block of tissue: GABA+, Glx, NAA, choline, creatine and mI+Gly, almost always as ratios to creatine.
In whomTransgender men (TM) before and during testosterone: 15 in Vienna, 28 retained in Ghent. One single-session study of 20 transgender women (TW).
What it can meanRegional shifts in tissue chemistry: hippocampal GABA+ relative to creatine in Vienna, parietal mI+Gly relative to creatine in Ghent. Not release, receptor function or cell counts.
What would settle itPaired counts for each endpoint, Vienna's participant overlap with the PET reports, a reconciled Ghent estimate, and any cohort followed through estradiol-based treatment.
The measurementWhat a spectrum records
Proton magnetic resonance spectroscopy (MRS) reads the chemistry of one block of tissue at a time. Hydrogen nuclei in each molecule resonate at their own characteristic frequency, so the result is a row of peaks rather than a picture.
The block is the voxel, usually a few cubic centimetres. Water dominates the raw signal and is suppressed, leaving metabolites present at millimolar levels. Position along the axis is chemical shift, given in parts per million (ppm). It is a property of the molecule rather than of the scanner, so N-acetylaspartate (NAA) appears at 2.01 ppm at any field strength. Peak area scales with how much of a molecule the voxel holds, summed over every cell type and compartment inside it.
Creatine, 3.03 ppm. The usual denominator. A ratio such as Cho/Cr moves if either peak moves.
GABA is here too, hidden. Its 3.01 ppm signal lies under creatine, and an unedited acquisition cannot separate it.
Glx. Glutamate and glutamine overlap at 3 tesla and are reported as one signal.
mI+Gly. Myo-inositol and glycine, reported together. Ghent's treatment finding is this peak divided by creatine.
NAA, 2.01 ppm. The tallest peak in most brain spectra.
Figure 2.1Two real spectra from the Ghent cohort. Single-voxel spectra at 3 tesla from the baseline report: (A) the left amygdala and anterior hippocampus voxel, (B) the right lateral parietal voxel. Peak labels and axes are the authors'; the numbered markers are ours. Collet 2021, Figure 1, right-hand spectra. CC BY 4.0. Cropped from the published figure with pixels unaltered; the participant MRI slices of the original panel are not reproduced.Representative spectra chosen by the authors. The original caption does not say which overlaid trace is the measured spectrum and which is the model fit.
Four cautionsA concentration is not release
Every number in this section describes a pool of molecules in a volume of tissue. Four properties of those pools limit what a change in them can mean.
1A pool, not a signal in transitMRS sees everything in the voxel: neuronal, glial, vesicular and metabolic. It cannot tell GABA released at synapses from GABA in storage or in metabolism. A lower GABA+ ratio is not, by itself, less inhibition.
2A ratio, not a concentrationGABA+/tCr falls if GABA+ falls or if creatine rises. In the Ghent baseline data, which groups differed depended on the ratio chosen. Ghent also reports water-referenced concentrations; Vienna reports ratios only.
3Mixed signalsGABA+ is GABA plus co-edited macromolecules. Glx is glutamate plus glutamine. Ghent's mI+Gly is myo-inositol plus glycine. A change in the sum does not say which part changed.
4Tissue and qualityA voxel holds grey matter, white matter and fluid in proportions that differ between people. Spectra that fail quality thresholds are dropped, so the analysed number shrinks region by region and visit by visit.
Two acquisitionsEdited and unedited spectra answer different questions
Ghent used an unedited sequence that captures the large peaks. Vienna and the Bangkok study used editing, which isolates GABA at the cost of almost everything else.
Unedited: PRESS
Point-resolved spectroscopy excites one box-shaped voxel and returns one spectrum: NAA, creatine, choline, mI+Gly and Glx. GABA stays hidden under creatine. Ghent used it at an echo time of 75 ms, with two voxels per session and tissue-corrected concentrations as well as ratios.
Edited: MEGA
Two interleaved acquisitions, with and without a narrow pulse at 1.9 ppm that perturbs only molecules coupled to that frequency. Subtracting them cancels creatine and leaves GABA+ near 3.0 ppm and Glx near 3.75 ppm. Vienna ran it as 3D spectroscopic imaging (MRSI); Bangkok used single voxels.
The scales differ too. Vienna acquired a grid of spectra across a large slab, nominal voxel about 4 cm³, and averaged them within regions segmented from each participant's own scan. Ghent placed two single boxes of 6 and 12 cm³. Their hippocampal values sample tissue differently.
Vienna · one slab, many voxels3D MEGA-edited MRSI
80 × 90 mm (× 80 mm deep) · z = -4 mm
Ghent · voxel 1single-voxel PRESS
20 × 15 mm (× 20 mm deep) · y = -7 mm
Ghent · voxel 2single-voxel PRESS
20 × 30 mm (× 20 mm high) · z = 34 mm
Figure 2.2Where the two cohorts measured, at one millimetre scale. Left: Vienna's 80 × 90 × 80 mm volume with its acquired 16 mm grid, over the five regions it analysed. Middle and right: Ghent's two single voxels. All three panels share one scale. MNI152 (2009) template and Harvard-Oxford atlas outlines, rendered with nilearn. Volume dimensions from Spurny-Dworak 2022 (Methods) and Collet 2021 (Methods).Box sizes are published; box positions are not, so they are placed here to cover the structures each paper names. Vienna's regions came from each participant's own segmentation, not from this atlas.
ViennaA GABA+ change inside the PET trial
Vienna's spectroscopy found one treatment effect, in hippocampal GABA+/tCr. It comes from the same registered trial as the newer harmine PET.
Spurny-Dworak 2022 scanned 15 TM and 15 cisgender women (CW) twice, at least 12 weeks apart. The actual interval was 146 ± 30 days for TM and 166 ± 40 days for CW. The TM were on seven different regimens: intramuscular or transdermal testosterone, in some cases with desogestrel or triptorelin.
In the hippocampus, GABA+/tCr fell in TM relative to CW across visits: a group-by-time interaction with p = .048 after Bonferroni correction across five regions, two ratios and two groups. There was no Glx/tCr effect in any region, no GABA+ effect in insula, putamen, pallidum or thalamus, and no correlation between hormone levels and either ratio.
15
peopleTM scanned at both visits
13
scansvalid hippocampal GABA+/tCr values at baseline
13
scansvalid values at follow-up
11–13
pairedTM with both values, by arithmetic; not reported
Same family as the PET
The report places itself within NCT02715232, the registration behind Kranz 2021 and Handschuh 2024. It therefore widens what that family measured, from MAO-A to GABA+ and Glx, rather than adding a verified independent cohort. Whether these 15 TM are the people in the PET reports cannot be read from the papers.
The authors read the result through hippocampal androgen metabolism and neurotrophic effects on interneurons, mechanisms drawn from animal work. The measurement cannot separate GABA from co-edited macromolecules, which they assume stable.
GhentThe one separately recruited cohort
Ghent is the only cohort here with molecular imaging before and during treatment that was recruited outside Vienna. Its treatment signal is a parietal decrease in mI+Gly/Cr; the amygdala voxel showed no simple change within TM.
Collet 2023 followed TM starting testosterone undecanoate, 1000 mg every 12 weeks, alongside cisgender men (CM) and CW with no intervention. Of 31 TM at baseline, 28 were retained at follow-up after 7.7 ± 3.5 months; the planned interval was six months, extended by COVID-19 delays. At follow-up, spectra of acceptable quality came from 25 TM in the parietal voxel and 20–22 in the amygdala voxel, depending on metabolite.
Within TM, parietal mI+Gly/Cr decreased (p = .015). Visit-by-group interactions appeared for parietal mI+Gly, mI+Gly/Cr and NAA/mI+Gly relative to CM, and for amygdala Glx/Cr relative to CW. In the amygdala/anterior hippocampus there was no within-TM change, although after treatment TM had lower Glx and Glx/Cr than CW.
One family, several reports
Collet 2021 is this cohort's baseline, and Kiyar 2022 combines fMRI with the same MRS data; neither adds people. Collet 2024 is a corrigendum that restores the ethics, funding and acknowledgement statements and changes no results. For counting: one family, one report on change during treatment.
One result needs reconciling before Ghent can enter any pooled estimate. The reported within-TM change in parietal mI+Gly/Cr is −0.045 (95% CI −0.076 to −0.015). The estimated means in the next sentence of the paper, 0.948 before and 0.944 during treatment, differ by only 0.004, which matches the paper's own later description of a 0.4% decrease. The supplementary medians, 0.954 and 0.908, imply a shift close to the coefficient.
Figure 2.3Three reported numbers for one change. Ghent's parietal mI+Gly/Cr in TM, as given in three places. Left: levels at baseline (T0) and follow-up (T1). Right: the change each source implies, beside the model coefficient. Collet 2023, Results and Supplementary material 4. Plotted directly from the reported values.Intervals are as reported: 95% CI for model estimates, interquartile range for medians. Different estimands, a scaling difference or a reporting error could each explain the gap; the paper does not say which.
Before pooling
Ask the authors which quantity the coefficient describes and on what scale, and how many TM had valid parietal spectra at both visits. The supplement's follow-up table is headed 31 TM, but its valid and missing counts sum to 28.
BangkokBackground, not change
The only spectroscopy report on TW found here is a single session. It shows that edited GABA measurement in TW is feasible, not how treatment changes it.
Jarukasemkit 2025 scanned 20 TW recruited for depressive symptoms, using MEGA-PRESS in the left hippocampus and dorsolateral prefrontal cortex. Lower hippocampal GABA/Cr went with higher depression scores until the participant with the highest GABA value was removed. About half were on GAHT (the methods say half, the results 45%); on- and off-treatment comparisons were not significant.
Table 2.1 · What each spectroscopy report measured. Valid counts are spectra passing quality control, from each report's tables or supplement; they are observations at one visit, not pairs.
Report
Cohort
Acquisition
Regions
Signals
Valid transgender spectra
Spurny-Dworak 2022
Vienna, NCT02715232
3D MEGA-LASER MRSI, 3 T, TE 68 ms
Hippocampus, insula, putamen, pallidum, thalamus
GABA+/tCr, Glx/tCr
TM 12–15 per region, ratio and visit
Collet 2021
Ghent, baseline
PRESS, 3 T, TE 75 ms
Left amygdala/anterior hippocampus; right lateral parietal
NAA, Cho, Cr, Glx, mI+Gly; ratios
TM 23 and 24, one visit
Kiyar 2022
Ghent, same MRS data
As Collet 2021, with fMRI
Amygdala/anterior hippocampus
Choline as moderator
TM 25, one visit
Collet 2023
Ghent, follow-up
PRESS, 3 T, TE 75 ms
As Collet 2021
As Collet 2021
TM 25 parietal, 20–22 amygdala at follow-up
Jarukasemkit 2025
Bangkok
MEGA-PRESS, 3 T
Left hippocampus, left DLPFC
GABA/Cr
TW 20, one visit
For a synthesisWhat spectroscopy buys, and what it does not
MRS widens a molecular synthesis in two ways: one more longitudinal family, and chemical signals that PET does not measure. It does not supply a larger paired sample.
What it adds
A second recruitment family followed through testosterone, from outside Vienna. Metabolites conventionally linked to inhibitory and excitatory chemistry, cell membranes and glia, alongside the PET targets. In Vienna, a second modality in the same trial, which could support within-person comparisons if participant overlap is confirmed.
What it does not
A common endpoint: the two cohorts share no metabolite ratio, region or sequence. A larger paired sample: 15 + 28 = 43 TM scanned twice is a nominal headcount. Any longitudinal data on estradiol-based treatment.
2
familieswith spectroscopy before and during treatment
1
familynot already in the PET evidence (Ghent)
43
peopleTM scanned twice across both cohorts, nominally
?
pairedTM with a valid pair for one shared endpoint: none exists
Hypotheses, not findings
Where spectroscopy might meet PET, three directions for a protocol:
If the Vienna spectroscopy and harmine participants overlap, hippocampal GABA+/tCr and regional MAO-A could be compared within the same people.
Vienna's baseline sex difference in insular GABA+/tCr, together with insular connectivity changes reported in TW (Reed 2023), makes the insula a candidate region for spectroscopy during estradiol-based treatment.
All longitudinal MRS found here concerns testosterone. A cohort of TW scanned before and during treatment is the clearest gap.
Section 03 places both spectroscopy families on the full cohort map, next to the PET families, and counts each participant once.
Deep dive: why GABA needs editing
GABA gives three groups of peaks, near 3.01, 2.28 and 1.89 ppm, and all three sit under larger signals: creatine, Glx and NAA. The 3.01 and 1.89 ppm groups belong to neighbouring hydrogens in the same molecule, so they are coupled. A narrow pulse at 1.9 ppm changes how the 3.01 ppm signal evolves, but only in molecules with that coupling.
Recording spectra with and without that pulse and subtracting them removes the uncoupled creatine peak and keeps GABA. Macromolecules with a resonance near 1.7 ppm are partly hit by the same pulse and contribute to the 3.0 ppm difference peak, hence the name GABA+. Glutamate and glutamine are co-edited and appear near 3.75 ppm. Separating GABA from macromolecules needs additional acquisitions, and Spurny-Dworak 2022 point to 7 T for separating glutamate from glutamine.
What this section does not show
That testosterone changes GABA release, synaptic inhibition or receptor function: MRS measures pools in a voxel.
A pooled effect across Vienna and Ghent: they differ in metabolites, regions, sequence, reference and interval.
How many TM contributed a valid before-and-after pair for any spectroscopy endpoint.
Any change during estradiol-based treatment: no longitudinal MRS in TW was found in this bounded search.
Anything about gender identity, clinical benefit or harm.
Ten PET and spectroscopy reports, three cohorts followed through treatment
Ten PET and spectroscopy reports trace back to five cohort families, and only three of those families measured the same people before and during treatment. Inside them, the number a change estimate rests on, people with a valid measurement at both times, is smaller than any headline sample size and is sometimes not reported at all.
Code-rendered · ten report panes over five cohort families, as in Figure 3.1 · Schematic motion
MeasuredWho contributed to each molecular report, counted in five separate units: reports, cohort families, scans, people and paired people.
In whomTen PET and MRS reports from five cohort families. Before-and-after molecular data come from two Vienna trials and one Ghent project.
What it can meanA synthesis has to count families, not papers, and paired people, not visits. Several paired counts cannot be read off the published reports.
What would settle itA per-endpoint flow diagram from each group, and confirmation of which people appear in which report.
UnitsFive ways to count the same study
A sample size means little until its unit is named. This literature mixes five units, and most overstatements come from adding one as if it were another.
□ ReportsPublications. Kranz 2014 and Kranz 2015 are two reports, and the 14 transgender women (TW) in the first reappear in the second.
◇ Cohort familiesRecruitment programmes, usually one trial registration or ethics approval. Reports from one family are not independent evidence.
○ Scans and visitsObservations. A person scanned three times contributes three of them.
● PeopleDistinct participants, counted once however many scans or reports include them.
●● Paired peoplePeople with a valid measurement of the same endpoint both before and during treatment. This is the denominator a within-person change actually rests on, and it can differ from one region or metabolite to the next.
The units nest, and each step narrows. A study enrols people, scans them at baseline, rescans a subset, loses some scans or spectra to quality control, and ends with a paired sample for each endpoint.
10
reportsPET and MRS reports: five PET, five MR spectroscopy (MRS)
5
familiescohort families behind those ten reports
3
familieswith before-and-after molecular data
43
peopletransgender men (TM) scanned twice across the two MRS cohorts, which is not one paired sample
Figure 3.1The cohort map: where each report's participants come from
Placed by what they measured and whom they recruited, the ten PET and spectroscopy reports fall into five families. The newer Vienna trial alone supplies three of them, plus two of the MRI reports used for context.
Figure 3.1Cohort map. Rows are what was measured; columns are cohort families, identified by trial registration or ethics approval. Each chip is one report, and its beads show the visits in the report's design (one bead: a single session). Brackets mark stated or near-certain reuse of the same participants. Hover over or tap a chip for its reference card; the buttons recount reports and families for each scope.
Methods and participant sections of each report, checked against the full text; Kranz 2018 at abstract level. Family assignments follow the registrations and reuse statements in the reports.There is no participant-level crosswalk between the Vienna trials, so their independence from one another is assumed, not shown. Which registration Kranz 2018 drew on was not checked.
Read down a column and the repetition shows. NCT02715232 is behind Kranz 2021, Handschuh 2024 and Spurny-Dworak 2022, and also behind Reed 2023; Konadu 2023 pools it with an earlier Vienna MRI trial, NCT01292785. The Ghent column holds three reports and one recruitment: Collet 2021 describes the baseline, Kiyar 2022 states that its spectroscopy data were published there before, and Collet 2023 follows the same project through testosterone treatment.
Read across a row and the thinness shows. Each PET target has exactly one longitudinal family: the serotonin transporter (SERT) in the earlier Vienna PET trial and monoamine oxidase A (MAO-A) in the newer one. GABA-edited spectroscopy has one longitudinal family (Vienna) and one single-session cohort (Bangkok), and the PRESS metabolite ratios have one family (Ghent). The Stockholm perfusion study and the Bangkok study measured each person in a single session, so neither can estimate change during treatment.
Five reports, one registrationNCT02715232 underlies two harmine PET analyses, the Vienna spectroscopy report, an insular connectivity report and half of a pooled hypothalamus analysis. That makes it valuable for linking modalities in the same people, and it is exactly why its reports cannot be read as replications of one another.
Film · about a minute · narratedFrom papers to people
The same arithmetic, animated on the three PET treatment reports: papers collapse into families, scans into people, and people narrow to pairs.
Film · 66 s · narrated Each participant is a thread through time, each visit a bead, and each paper a window onto the threads. The counts are the published ones. Which thread carries which bead is schematic, because no report publishes participant-level allocation.
Kranz 2015, Methods and Table 1; Kranz 2021, Results and Table 2; Handschuh 2024, Results and Table 1.The eight-to-ten range is our arithmetic bound, not a reported count.
Those three moves are the method of this section, and every number below carries its unit so that a reader can make them too.
Table 3.1The sample ledger: four numbers per report, not one
For each longitudinal molecular report, the ledger keeps enrolled, baseline-valid, follow-up-valid and explicitly paired counts in separate columns. Where a paper does not report a number, the cell stays open.
The column that matters most for a change estimate is the last, and few rows fill it from an explicit statement. Handschuh 2024 states that 17 participants completed both PET scans and breaks them down by group. In Kranz 2015 everyone rescanned had a baseline scan by design. Elsewhere the paired count has to be bounded by arithmetic, which narrows it but does not settle it.
Transgender participants only; controls omitted. "By design": every rescanned participant had a baseline scan. "By arithmetic": a bound from the number started and the valid counts at each visit (our inference, not a reported count). Dots are counts, not individuals. Sources: Methods, tables and supplements of each report, checked against the full text.
Report · endpoint
Family
Started
Valid at baseline
Valid at follow-up
Explicitly paired
Kranz 2015 SERT · 4 weeks
Vienna SERT
3314 TM + 19 TW
33
259 TM + 16 TW
25by design
Kranz 2015 SERT · 4 months
Vienna SERT
same 33
same 33
2411 TM + 13 TW
24by design
Kranz 2015 · people with any follow-up scan
The two follow-up visits overlap; they are not added.
not reported27–32 by arithmetic
Kranz 2021 MAO-A · ~4 months
Vienna multimodal
1811 TM + 7 TW
1610 TM + 6 TW
106 TM + 4 TW
not reported8–10 by arithmetic
Handschuh 2024 MAO-A · MRI-defined regions
Vienna multimodal
18same PET subgroup
1811 TM + 7 TW
116 TM + 5 TW
11stated; not added to 2021
Spurny-Dworak 2022 hippocampal GABA+/tCr
Vienna multimodal
15TM
13
13
not reported11–13 by arithmetic
Collet 2023 parietal metabolite ratios
Ghent
31TM; 28 retained
not reported in this report
25of 28 retained
not reportedat most 25
Collet 2023 amygdala / anterior hippocampus
Ghent
same 31
not reported in this report
20–22by metabolite
not reportedat most 20–22
One report can also carry several paired samples: valid values in Spurny-Dworak 2022 range from 12 to 15 per visit across regions and ratios. The reports' mixed models can use unpaired scans, so each analysis is sound on its own terms; a synthesis of within-person change still needs the pairs.
Counting trapsSix places where the arithmetic goes wrong
Each trap is a sum that looks natural and is wrong. All six come from the published reports themselves. None needs an error by the authors, only a reader who adds across units.
1
Adding visits. The DASB study measured 25 transgender participants at four weeks and 24 at four months: 49 follow-up scans. Largely the same people came back, so far fewer than 49 people had any follow-up. The union is not reported. Arithmetic puts it between 27 and 32: at least the larger visit in each group (11 TM, 16 TW), at most everyone who could have returned (14 TM, and 18 TW because one participant moved away after baseline).
2
Reading follow-up observations as pairs. Kranz 2021 reports 10 usable four-month harmine scans in transgender participants. A pair also needs a usable baseline from the same person. With 16 usable baselines among 18 people assessed, between 8 and 10 of those ten can be pairs. Ten follow-up observations are not a verified ten pairs.
3
Adding two analyses of one trial. Handschuh 2024 reports 11 transgender participants who completed both harmine scans. The trial, the PET subgroup sizes (11 TM, 7 TW, 9 cisgender women (CW), 8 cisgender men (CM)) and the scanner failure that ended PET acquisition are those of Kranz 2021. Adding 10 and 11 would count largely the same people twice. The two papers also divide the same 52 usable scans differently (Figure 3.2A), so even inside the family a pooled unique total stays unresolved.
ASame trial, same 52 usable scans, two allocations
BOne Ghent baseline, three transgender-men denominators
Collet 2021
29TM in the baseline report23 / 24TM with valid spectra (amygdala / parietal)
Kiyar 2022
30TM recruited25TM in the MRS analysis
Collet 2023
31TM at baseline28TM retained at follow-up25TM with valid parietal spectra at follow-up
Figure 3.2Same family, different denominators. (A) The two harmine reports from NCT02715232 each report 52 usable scans, split 33 baseline + 19 follow-up in Kranz 2021 and 35 + 17 in Handschuh 2024. (B) The Ghent reports give three different baseline headcounts for one recruitment.
Kranz 2021, Results and Table 2; Handschuh 2024, Results and Table 1; Collet 2021, Results; Kiyar 2022, Methods; Collet 2023, Results and Table 1.The differences may reflect analysis-specific inclusion rules rather than errors; only the authors can reconcile them.
4
Pooling the MRS cohorts into one paired sample. Fifteen TM were scanned twice in Vienna and 28 were retained at follow-up in Ghent: 43 nominal repeat-scanned participants in two independent cohorts. That is not a common valid paired-metabolite sample. The cohorts used different sequences (edited 3D MRSI of GABA+ and Glx versus single-voxel PRESS ratios) in different regions, and their valid counts are smaller and endpoint-specific: 13 at each visit for Vienna hippocampal GABA+/tCr, 25 at follow-up for Ghent parietal ratios.
5
Counting the Ghent baseline again. Collet 2021 describes the baseline, Kiyar 2022 re-analyses its spectroscopy with fMRI, and Collet 2023 reports the follow-up. That is three reports from 1 recruitment, with three baseline headcounts (29, 30 recruited, 31; Figure 3.2B). Neither cross-sectional report adds people to the longitudinal one.
6
Trusting a table header. The availability table in Collet 2023's supplement gives the follow-up total as 31 TM, which is the baseline number. In the same table, valid plus missing spectra sum to 28 for every metabolite, matching the 28 retained in the main text. The 28 is the better-supported figure, but the table cannot say which of those people also have a valid baseline spectrum, so complete-case pairs stay open.
Why the bounds stay boundsThe ranges here (27–32, 8–10, 11–13) are arithmetic consequences of the published tables. They limit what the true counts can be. They are not estimates, and no point within them is more likely than another.
AsksWhat would settle the counts
None of these gaps needs new data. Each could be closed with information the research groups already hold, and each ask is specific enough to answer in a table.
A flow diagram per endpoint for each longitudinal report: enrolled, scanned at baseline, valid at baseline, rescanned, valid at follow-up and paired, for every region and ratio analysed.
Participant allocation within each family. Which people appear in Kranz 2021 and in Handschuh 2024, and why the same 52 scans divide 33 + 19 in one and 35 + 17 in the other. Likewise for the Ghent participants across Collet 2021, Kiyar 2022 and Collet 2023.
A crosswalk between the Vienna registrations. Whether anyone took part in more than one of NCT01065220, NCT01292785 and NCT02715232, and which trial Kranz 2018 drew on.
Protocol information. Planned visits and intervals, quality-control thresholds, and the status of planned but unreported measurements such as the [11C]AMT arm registered under NCT02715232.
Paired summaries. Mean and standard deviation of within-person change, or the baseline-to-follow-up correlation, for each endpoint. A synthesis needs these, and mixed-model outputs do not provide them.
What the answers would buyWith these, the ledger could carry one unique-person total per family and one paired count per endpoint. That is a precondition for any quantitative synthesis, and it would let reports from one trial be read as complementary views of the same people rather than counted twice.
ContextSelected MRI: mechanism, not more cohorts
Four MRI analyses help interpret the molecular signals. None adds a molecular study, and three of them reuse the Vienna multimodal trial already counted.
Handschuh 2024 is the closest bridge: grey-matter and microstructure change measured in the same people whose MAO-A was scanned. Reed 2023 analyses insular connectivity in 14 TM and 11 TW scanned twice in the newer Vienna trial. Konadu 2023 pools that trial with an earlier Vienna MRI trial to study hypothalamic volume in 38 TM and 15 TW scanned twice. Kranz 2018 reports hypothalamic diffusion in 25 TM over the same three-visit schedule as the DASB study; its registration and overlap were read only at abstract level.
Their samples are larger than the molecular ones, which makes them tempting to add. They should not be added. A Konadu participant from NCT02715232 may also be a harmine, spectroscopy or connectivity participant, and the report does not list which trial each person came from. The right use of these studies is mechanistic: what changed in tissue alongside a molecular signal, in people already counted.
What this section does not show
That the Vienna trials share participants. It shows only that their independence has not been demonstrated.
Whether any finding is right. Sections 01 and 02 cover what was measured and what it can mean.
True paired counts. The ranges given are arithmetic bounds, not reported numbers.
A complete eligibility count. The ten reports are source-checked candidates from a bounded search, not the output of a full systematic review.
A focused paper would add a careful count, not a longer list
The closest reviews already read several of the same primary reports, each under its own question. A focused synthesis of PET and MR spectroscopy during hormone therapy, with selected MRI as context, could complement them by counting reports, cohorts and valid pairs once, and by reading each measure for what it measures. How much it adds depends on the current study table of the scoping review with Michael Winterdahl, which we have not yet seen.
Code-rendered (WebGL) · three review questions as lenses over the same study threads · Schematic motion
MeasuredHow three existing reviews cover the primary reports: two checked against the full text, one read from its published abstract.
In whomThe same reports this report audits: five PET, five MRS and three further MRI reports, from a handful of cohort families.
What it can meanA focused PET + MRS critical or scoping synthesis, with baseline contrasts kept apart and no pooling per target.
What would settle itThe scoping review's current study table, reconciled harmine pairs, and the status of a planned third tracer.
LandscapeThree reviews, three different questions
The closest reviews are not rivals to a new paper. They ask different questions of overlapping material, and each does what it set out to do.
Positron emission tomography (PET) and magnetic resonance spectroscopy (MRS) during gender-affirming hormone therapy (GAHT) have been touched by three reviews that matter here. Two were read in full; the third, the most recent and the closest in scope, is available to us only as a congress abstract.
Narrative overview · 2020Kranz 2020
Asks: what longitudinal MRI and PET studies show about the brain during GAHT.
Molecular content:1 PET treatment report, the 2015 serotonin transporter (SERT) study with [11C]DASB, described across baseline, one month and four months.
Written before the monoamine oxidase A (MAO-A) and MRS treatment reports appeared.
Systematic review · 2021Frigerio 2021
Asks: how brains differ by gender identity or sexual orientation before treatment; post-treatment data are excluded by design.
Molecular content:3 PET reports (Berglund 2008, Kranz 2014 and the baseline of Kranz 2015) and 1 single-photon emission computed tomography (SPECT) report, too few to meta-analyse.
Search to January 2018 for gender identity; a Medline update to March 2021.
Scoping review · ESSM 2025 abstractNielsen 2025
Asks: how the brain changes during GAHT across imaging methods, and how future studies could be designed better.
Content:52 included studies from 212 records found in PubMed and Embase.
The accessible abstract has no study table, search cutoff, PET-target count or cohort map. The current manuscript was not available to us.
Read side by side, they show that overlap belongs to a question, not to a paper. The 2015 DASB study appears in Kranz 2020 as a treatment study and in Frigerio 2021 through its baseline comparisons alone. Whether a new paper overlaps an existing review therefore has to be judged report by report and question by question.
Complementary, not competingThe scoping review maps GAHT imaging broadly, across methods and regions. A molecular synthesis would go narrow and deep: two PET targets, a few MRS metabolites, and an exact account of who was measured twice. The two can be read together, and could be built together.
Figure 4.1Which review covers which report
Of 13 primary reports, Kranz 2020 reviews two and Frigerio 2021 includes four, two of them only through baseline data. Everything published after 2020 is new to both, and the scoping review's coverage stays unknown until its study table can be compared.
Look through
Primary report
Kranz 2020narrative · longitudinal
Frigerio 2021systematic · baseline only
Nielsen 2025scoping · abstract
ProposedPET + MRS synthesis
PET
Berglund 2008Stockholm · perfusion PET, untreated
Not included
Included
Unknown: no study table available
Background: baseline or cross-sectional
Kranz 2014Vienna SERT trial · DASB, baseline
Cited as a pointer, not reviewed
Included
Unknown: no study table available
Background: baseline or cross-sectional
Kranz 2015Vienna SERT trial · DASB, treatment
Included
Baseline data only
Unknown: no study table available
Core: change during treatment
Kranz 2021newer Vienna trial · harmine
Published after the review
Published after the review
Unknown: no study table available
Core: change during treatment
Handschuh 2024newer Vienna trial · harmine + MRI
Published after the review
Published after the review
Unknown: no study table available
Core: change during treatment It is also the direct MRI–molecular bridge.
MRS
Collet 2021Ghent · baseline
Published after the review
Published after the review
Unknown: no study table available
Background: baseline or cross-sectional
Kiyar 2022Ghent · fMRI + MRS, baseline
Published after the review
Published after the review
Unknown: no study table available
Background: baseline or cross-sectional
Spurny-Dworak 2022newer Vienna trial · GABA+, Glx
Published after the review
Published after the review
Unknown: no study table available
Core: change during treatment
Collet 2023Ghent · treatment follow-up
Published after the review
Published after the review
Unknown: no study table available
Core: change during treatment
Jarukasemkit 2025Bangkok · cross-sectional TW
Published after the review
Published after the review
Unknown: no study table available
Background: baseline or cross-sectional
Selected MRI
Kranz 2018older Vienna programme · diffusion
Included
Baseline data only
Unknown: no study table available
Context: selected MRI
Konadu 2023two Vienna trials pooled · volume
Published after the review
Published after the review
Unknown: no study table available
Context: selected MRI
Reed 2023newer Vienna trial · connectivity
Published after the review
Published after the review
Unknown: no study table available
Context: selected MRI
IncludedBaseline data onlyCited as a pointer, not reviewedNot includedPublished after the reviewUnknown: no study table availableCore: change during treatmentBackground: baseline or cross-sectionalContext: selected MRI
Figure 4.1Review coverage, report by report. Rows are the primary PET, MRS and selected MRI reports discussed in this report; columns are three existing reviews and the proposed synthesis. A filled cell means the review includes the report; half-filled means only its baseline data are used. Choose a review above to look through it alone; hover a cell for its reading.
Kranz 2020, full text (introduction, molecular section, Table 1); Frigerio 2021, full text (Methods, Table 1, metabolic section, Table 3); Nielsen 2025, publisher abstract.Unknown is not absent: the scoping review may include many of these reports. "Published after" means after the review was written or after its search window. The proposed column is a working plan, not a protocol.
Two things stand out. The MAO-A reports, both longitudinal MRS cohorts and the newer Vienna MRI analyses all postdate the two full-text reviews, so a molecular update has real material to add. And the one review that could already cover that material is the one whose table we have not seen. Any statement about what a new paper adds should wait for that comparison, and no paper here should claim to be the first review of anything.
OptionsThree ways to write it
A focused PET + MRS synthesis is the working preference. A PET-only paper is possible but thin, and an all-modality catalogue would mostly re-cover ground that existing reviews, including the scoping review (Nielsen 2025), already map.
Option 1Working preferenceFocused PET + MRS synthesiscritical or scoping; selected MRI as context
PET change
MRS change
Baseline, in a separate section
Selected MRI
Other MRI, fMRI
Would claim
What PET and MRS have measured in the brain during GAHT, in which cohorts, on how many valid pairs, and what each measure can and cannot mean.
Rests on
5 reports with before-and-after molecular data, from 3 cohort families: Vienna SERT, newer Vienna, Ghent.
Needs
Per-endpoint pair counts; who overlaps across Vienna's PET, MRS and MRI reports; the scoping review's table; a written protocol if it is to be called scoping or systematic.
Main risk
A small base. All longitudinal MRS is in transgender men (TM), and one Vienna trial supplies three of the five core reports.
Overlap with existing reviewsmoderate, pending the table
Option 2PET-only critical papershort; two targets
PET change
MRS change
Baseline, in a separate section
Selected MRI
Other MRI, fMRI
Would claim
What SERT and MAO-A binding changes during GAHT do and do not show.
Rests on
3 treatment reports from 2 Vienna families, on two targets.
Needs
Reconciled harmine allocations and pairs; a settled answer on the planned third tracer.
Main risk
Thin: two targets, each from one Vienna trial; best suited to a short critical paper.
Overlap with existing reviewslow to moderate
Option 3All-modality cataloguestructure, function and molecules
PET change
MRS change
Baseline
Selected MRI
Other MRI, fMRI
Would claim
What brain imaging of every kind has found during GAHT.
Rests on
Dozens of MRI and functional MRI (fMRI) reports; the scoping review alone includes 52 studies.
Needs
A full screen of 573 retrievals and more databases, then extraction and cohort mapping across many MRI studies.
Main risk
Heavy overlap with existing reviews, including the scoping review; much effort for a modest addition.
Overlap with existing reviewshigh
Figure 4.2Three scopes on the same evidence. For each option: what the paper would claim, what it rests on (with units), what it still needs and its main risk. The scope strip shows what each option includes (solid), keeps in a separate role (half) or leaves out (open).
Counts from the PET and MRS chapters of this report; scoping-review count from Nielsen 2025.The overlap gauges are our judgement, not a measured overlap; option 1's gauge depends on the scoping review's study table.
The preferred design is a critical review with a reproducible search, or a scoping review if the team prefers a formal map. Its core is change during treatment: three PET reports from two Vienna families and two longitudinal MRS reports, one of them from the separate Ghent cohort. Baseline group contrasts sit in their own background section: Berglund 2008, Kranz 2014, Collet 2021, Kiyar 2022 and the cross-sectional Bangkok study in transgender women (TW), Jarukasemkit 2025. That way a difference before treatment is never read as an effect of treatment.
There would be no meta-analysis per target. Two families and two PET targets leave nothing independent to pool, and the two harmine analyses ask different regional questions of one study family, so their different results are not a replication test. The contribution would be the accounting itself: baseline versus treatment comparisons, cohort reuse, valid paired endpoints, and what binding potential, distribution volume and metabolite ratios each measure.
A joint route, for the readers to decideBecause the scoping review with Michael Winterdahl already maps GAHT imaging across methods, one natural route would be to build the molecular synthesis with that team, for example as a focused molecular companion to the broad map, or as a deeper molecular layer within it. This is an option to weigh, not a plan; it depends on where that manuscript stands and on what Annie and Michael want.
Selected MRIMRI explains the molecules; it never adds cohorts
Four MRI analyses earn a place as context. Three reuse the newer Vienna trial already counted; the fourth, Kranz 2018, comes from the older Vienna programme, and its overlap with the DASB cohort is unresolved. None is counted as more molecular evidence or more people.
Handschuh 2024 is the direct bridge: grey-matter and microstructure changes measured in the same people and visits as harmine PET, with no significant MAO-A change in the MRI-defined regions. Its 11 transgender participants with both PET scans are already in the PET core.
Konadu 2023 follows hypothalamic volume in 38 TM and 15 TW scanned twice, pooled from two Vienna trials: the newer multimodal trial NCT02715232 and an earlier MRI trial, NCT01292785. It is a reanalysis of earlier datasets, not a third cohort.
Reed 2023 reports insular connectivity in 14 TM and 11 TW from the newer Vienna trial, a useful lead next to the insula MRS measurements.
Kranz 2018 follows hypothalamic diffusion in 25 TM over four months, from the older Vienna programme; read at abstract level here.
Used this way, MRI helps interpret a molecular result (did tissue change where binding did not?) without being counted as more molecular evidence or more independent people.
Before committingSeven open items, owners to be agreed
None of these needs new data. Most need a careful re-reading or a short, collegial question to the groups that ran the studies.
Table 4.1 · Open items before choosing a scope. The owner column is deliberately empty: who takes what is for the readers to agree.
#
Open item
Why it matters
Possible route
Owner
1
Compare against the scoping review's current study table
Settles overlap and any statement of what a new paper adds
Ask for the current manuscript version or its table
to agree
2
Reconcile harmine sample allocations and exact valid pairs
The 2021 and 2024 papers divide 52 usable scans differently; 10 four-month observations are not 10 verified pairs
Re-extract both tables; possibly ask the Vienna authors
to agree
3
Settle the planned [11C]AMT arm
Registered and funded, with no verified transgender output
Ask the Vienna group whether AMT scans were acquired
to agree
4
Pair counts per MRS endpoint
15 Vienna and 28 Ghent TM were rescanned; valid pairs are fewer and differ by region
Supplements first, then an author query
to agree
5
Full screen of the retrievals, if a systematic review is wanted
573 records are retrievals, not eligible studies
Two-reviewer screen; add Embase, Scopus, Web of Science and registries
to agree
6
Separate adult GAHT, puberty suppression and regimens
Exposures differ; estradiol plus antiandrogen is not estradiol alone
Write eligibility rules before extraction
to agree
7
Targets not found here: steroid receptors, SV2A, dopamine, neuroinflammation
A qualified search finding, not proof of absence
A targeted search before any gap is stated
to agree
The third item deserves care in both directions. The trial registration NCT02715232 and the Austrian Science Fund programme FWF KLI504 planned serotonin-synthesis PET with [11C]AMT (α-methyl-L-tryptophan) alongside harmine, before and after about four months of treatment, and the EU register entry (EudraCT number 2015-000502-19) marks the trial completed. No transgender AMT scan count or result was found. The harmine papers report that the PET scanner was irreparably damaged in December 2019. That explains the truncated harmine data; it does not show that every planned AMT scan was cancelled.
The seventh item is phrased narrowly on purpose. This pilot did not find GAHT studies of steroid receptors, synaptic vesicle glycoprotein 2A (SV2A), dopamine or neuroinflammation targets, but its search was bounded. That is a reason to search for them deliberately, not a gap to announce.
GuardrailsWhat not to claim
Each line is a sentence that would be easy to write and wrong to publish. Together they are the spine of any paper built from this pilot.
Not "the first review" of molecular imaging during GAHT, or any novelty claim before the comparison with the scoping review.
Not three independent longitudinal PET cohorts: three treatment reports come from two families.
Not49 DASB participants: 25 at four weeks and 24 at four months are largely the same people.
Not ten harmine pairs from ten follow-up observations, and no pooled unique harmine total while the two papers' allocations differ.
Not43 MRS pairs: 15 Vienna and 28 Ghent TM were rescanned, and valid pairs per endpoint are fewer.
Not a failed replication: the 2024 harmine analysis asks a different regional question of the same study family.
Not serotonin from binding, or concentration from a ratio; and nothing about clinical benefit, harm, cell loss or gender identity.
Not an estrogen effect from estradiol plus antiandrogen, and not a treatment effect from the cross-sectional Bangkok study.
Not a systematic review: 573 records were retrieved, not fully screened.
Not AMT as a measured target, nor as a cancelled one.
Not an absence of steroid-receptor, SV2A, dopamine or neuroinflammation studies.
What this section does not show
What the scoping review with Michael Winterdahl includes. Its study table was not available, so its column is unknown, not empty.
That any option is novel. That waits for the comparison in item 1.
A systematic search: only PubMed and Europe PMC were queried programmatically, and the 573 retrievals were not fully screened.
A measured overlap: the gauges in Figure 4.2 are our judgement.
Any clinical recommendation, or a case for collecting new data now.
Each card states the design, who was measured, what the source showed, its main limitation, and how deeply it was read. Filter by kind, or search.
GlossaryTerms, defined once
The same definitions that appear on hover, gathered in one place.
MethodHow this synthesis was made
This is an exploratory pilot. It is not a completed systematic review, and its counts are source-checked candidates rather than a final eligibility denominator.
Reproducible PubMed and Europe PMC queries were run for PET, MR spectroscopy, longitudinal MRI and existing reviews, with a modality-blind rescue search and citation chasing. Together they returned 573 de-duplicated records. These are retrievals and have not been fully screened. Every primary report cited here was read in full, including methods, tables and supplements where available, unless its card says otherwise. Sample counts were extracted separately for enrolment, valid baseline, valid follow-up and explicit pairs. Cohort families were assigned from trial registrations, ethics numbers, funding and authors' own statements of reuse; participant-level identity across programmes has not been established. Registry and grant records were consulted for planned tracer arms. Unresolved values are left unresolved.
The published version of the existing scoping review (Nielsen 2025) is a congress abstract; its study table was not available, so overlap is assessed against the abstract and two earlier full-text reviews (Kranz 2020, Frigerio 2021).
Counted Once · Landau Lab, Aarhus University · 3 October 2026 · Illustrations are labelled with the real material they were made from. Films are code-rendered; anything simulated is marked “Schematic”.