MRI tutorial

MRI, start to finish.

Three Siemens DICOM folders that are not three subjects. This is the dataset the general walkthrough uses, chosen because every datatype in the workflow appears in it at least once: anatomical, functional, diffusion, fieldmaps, and the physiological signals the scanner stored inside the DICOM files.

This page builds on the general walkthrough.

The GUI walkthrough explains each control once, for every modality. This page covers what is specific to MRI, and what this dataset in particular is there to teach.

Example data

Oldenburg neuroimaging unit (3 folders, 2 subjects)

Three source folders. The first two, OL_0001 and OL_0002, share a patient identifier and collapse into sub-001 with ses-pre and ses-post. The third, OL_0003, is a different person and stands alone as sub-002 with no session. 394 DICOM files, 3T Prisma.

Download MRI sample dataset →

The data

Three folders, two subjects.

▾ neuroimaging_unit/ 394 DICOMs
▾ OL_0001/ sub-001 / ses-pre · 45 DICOMs, 9 series
T1w, two functional tasks (sparse, rest), two fieldmap pairs
▾ OL_0002/ sub-001 / ses-post · 62 DICOMs, 10 series
Multiband BOLD, one fieldmap pair, one physiological log
▾ OL_0003/ sub-002 · 287 DICOMs, 17 series
T2w, three functional tasks with reference volumes, diffusion with a reversed-direction b0, two physiological logs

What this dataset is here to teach

Folder names are not subject identities. Somebody named these folders in the order the scans happened, and a tool that trusts folder names would produce three subjects from two people. That is not a hypothetical: it is the single most common way an MRI dataset comes out wrong, and it is silent, because three subjects is a perfectly valid BIDS dataset.

How the grouping works

Subjects come from the DICOM header, not the folder.

The scan reads each folder's DICOM files and groups them by the patient identifier and patient name the scanner wrote, which is a fact about the person rather than a fact about how the export was organised.

  • OL_0001 and OL_0002 share a patient identifier, so they are one subject. The study dates order them, and the earlier becomes ses-pre, the later ses-post.
  • OL_0003 is a different identifier, so it is a second subject. It has one visit, so no session entity is written: a session number that only ever takes one value adds nothing and clutters every filename.

Sessions come from the study identifier and date in the same way. This is the MRI counterpart of the path heuristics EEG and MEG have to fall back on, and it is more reliable, because it is reading what the scanner recorded rather than what somebody typed.

Check the id column anyway.

Grouping is only as good as the identifiers. A site that reuses a patient identifier between people, or anonymises everyone to the same value, will produce one subject where there should be many. The column is editable for exactly this reason.

Step 1

Scan the folder.

Create a project, point Raw input at the top of the tree, and press Scan. Every DICOM file is opened and its header read: the acquisition parameters, the study and series identifiers, the patient identifiers used to group subjects, the timestamps, and whether there is image data inside at all.

You get 33 rows. Twenty-one are keepers and twelve are skipped automatically: the localisers the scanner takes before the real acquisitions, and the report objects it writes alongside them. Both are recognised and set aside rather than deleted, and any of them can be re-included with a click.

The scan on this dataset. The status line names the stage it is in, and the chips settle on 21 keepers and 12 skipped, 33 rows in total.

Leave probe convert on

With probe convert enabled, the scan additionally runs the converter once per series and reads the sidecar it produced. That is how the tool knows, before you have typed anything, which metadata fields the conversion will answer by itself. Those fields are then shown to you as already answered rather than asked for. On this dataset it adds around twenty-five seconds and saves a great deal of typing.

Step 2

Read the inventory.

One row per series. Before changing anything, read what the scan decided, and pay attention to three things in particular.

The inventory. Rows are tinted by status, and the sequence column keeps the original scanner label beside the predicted BIDS name, which is how you check a classification without opening anything.
Inventory rows the scan set aside, each carrying its reason
The twelve skipped rows account for themselves. The teal row is a scanner report with no image data inside it, the dimmed ones are localisers, and the reason for each is in the issues column. None of them are deleted: untick to exclude, tick to bring back.

The conf column

Each classification carries a confidence score. A high value means the converter's own guess and the schema agreed. A lower one means the decision came from a sequence-name rule, which is a good guess about a naming convention rather than a fact. Sort by confidence and read the bottom of the list: that is where a misclassification will be if there is one.

The sequence column, beside the predicted name

The original scanner label sits beside the BIDS name the row will produce. Reading the pair is the check. ses-pre_task-rest_bold next to a sequence called rest is right; task-rest next to a sequence called localizer is not.

Rows the scan set aside

A dimmed row is excluded. A teal row with a crossed-circle badge is something with no image data inside it at all, a Siemens TENSOR map or a PhoenixZIPReport, which cannot be converted by anything and is excluded with the reason shown. That distinction matters: without it, an unconvertible object looks exactly like a failed conversion of something that mattered.

Step 3

Curate.

The decisions on an MRI dataset are mostly about what to leave out and what to call things.

Exclude what should not be in the dataset

The scan already excluded what it could recognise. What it cannot know is that the second T1w was the one worth keeping because the subject moved during the first, or that a run was aborted and repeated. Untick the include box on those.

Give the tasks real names

Task labels come from the sequence names, which means they inherit whatever the protocol was called on the console. Check them for consistency above all: task-Rest in one row and task-rest in another produces two tasks in a dataset that has one, and no validator will object.

Bulk edit. Select the rows, choose the column, type the value once. The fastest way to fix a naming inconsistency that runs across a whole session.

Teach the scanner your site's conventions

If your protocol names something in a way the classifier does not recognise, and it will happen on your own data even though it does not happen here, add a rule in Settings rather than editing rows every time you scan. A hint forces any series matching a pattern to a datatype, suffix and task you choose. An exclusion drops anything matching a name or path. Both are constrained by the schema, saved with the project, and read by the command line through the same file.

Scan rules. A hint routes anything matching a pattern to the datatype and suffix you pick; an exclusion drops it. Written once, applied to every scan of that protocol thereafter.
Step 4

The metadata, and how little of it MRI needs.

MRI is the modality where the file answers the most. Echo time, repetition time, flip angle, field strength, coil, phase encoding direction, slice timing: all of it is in the DICOM header, and the conversion reads it. This is why there is no MRI-specific metadata block to fill in, and why the template's MRI sections are mostly the green already-answered kind.

What remains is about the study rather than the acquisition:

  • The dataset fields. Name, authors, licence, acknowledgements, funding. Nothing in a DICOM file knows any of these, and a dataset without authors is the most common thing wrong with published BIDS data.
  • Task descriptions. What the participant was asked to do, and the instructions they were given. Recommended by the standard and impossible to derive.
  • Participant information. Sex, age and handedness reach participants.tsv. Age and sex are read from the DICOM header where present; handedness never is, and is never guessed.
The metadata template on an MRI dataset, where most acquisition fields are already answered by the conversion
Mostly already answered. For MRI the work is at the top of this dialog, in the questions every modality shares, and in the task descriptions. The acquisition parameters sit folded away in the green Already answered by the conversion blocks, shown so you can check them rather than type them again.
Step 5

Convert.

Open the BIDS preview in the bottom panel and read the tree it will produce. Then press Run conversion. Image series go to dcm2niix; the physiological logs go to bidsphysio. You do not choose: the row's content does.

On this dataset: 24 NIfTI files across anatomical, functional, diffusion and fieldmap, three physiological tables, three .bval and .bvec pairs beside the diffusion series, and 29 sidecars. No failures.

The conversion running. Each subject is built in a private staging folder and moved into the dataset only once it has succeeded, so an interrupted run leaves nothing half-written.

Removing the face during the conversion

A T1-weighted head scan contains a face, and a face can be rendered from one. Settings → Convert → Deface anatomical and PET images removes it as each image is written, so the identifiable version never exists inside the dataset at all. It costs about two seconds per image, it touches only anat/ and pet/, and it can equally be done afterwards from the Editor if you would rather look at the images first. Defacing and skull stripping →

Residual outputs, dropped by default

dcm2niix sometimes splits one series into the real image plus derived single-volume copies. They have names like ..._bolda or ..._Eq_, and they are almost never wanted: left in place they double the apparent number of runs. They are dropped by default, and a setting keeps them if your workflow needs them. Fieldmap, complex-valued and multi-output diffusion series are not affected.

Step 6

What happens after the files are written.

Two MRI-specific repairs run automatically once the conversion is done, and both fix things that are valid BIDS and useless.

Fieldmaps that point at something

A fieldmap without IntendedFor is structurally valid and no pipeline can use it, because nothing says which images it is meant to correct. The enrichment fills it in from what the scan measured: which functional and diffusion runs sit in the same session with matching acquisition parameters.

The reversed-direction volume becomes a fieldmap

On sub-002 there is a single-volume diffusion acquisition with reversed phase encoding. It is not diffusion data in any useful sense: it exists to serve as the reference for distortion correction. The classifier recognises it and routes it to the fieldmap folder with the right suffix, so the pipeline that needs it can find it. No configuration file, and nothing for you to do.

The physiological recordings inside the DICOM files

Siemens scanners record respiration, cardiac and pulse-oximetry signals alongside the images and store them inside the DICOM data. Those become BIDS physiological files, compressed tables with their own sidecars, beside the functional run they were recorded with. They are easy to miss entirely, because nothing in the folder looks like a physiological recording.

Looking at one, in the Editor

A *_physio.tsv.gz opens as a table, which answers nothing about the shape of a trace. Above the table is a Visualize button, and it opens the same viewer MEG and EEG use, with the same filters, the same spectrum and the same controls described in the EEG tutorial.

The button is offered on the strength of the sidecar, not the filename. A table with a sampling frequency is a continuous recording whatever it is called; a table of onsets, like _events.tsv, is not, and is not offered a viewer it would have nothing to draw in.

A single physiological channel in the shared signal viewer, with the channel-selection controls hidden
One channel, so the channel controls are gone. Which of one channel to show, and how many of them, are both questions with one answer, so the type filter, the visible-count box and the channel picker are hidden. Fit all puts the whole recording in the window in one click, instead of winding the window length up a step at a time.

All of this run, together

BIDS splits a run's physiological data by the recording entity, so the cardiac trace, the respiratory belt and the trigger are three files describing one acquisition. Reading them one at a time is reading a three-channel recording one channel at a time, and whether the trigger lines up with the belt is exactly the question people bring to them.

All of this run loads the lot. They are put on the fastest one's grid, because that is the only rate at which nothing is thrown away, each at its own StartTime offset. Upsampling is nearest-sample rather than interpolated, and a stretch a slower or shorter recording does not cover is drawn as a gap rather than filled with an invented value.

Four physiological recordings of one run drawn together on a single time axis
Four recordings of one run, on one axis. From the top: the ECG, the cardiac trace, the respiratory belt and the trigger, each sampled at its own rate in its own file. With more than one channel the type filter and the visible-count control come back, because now they have something to choose between.
A missing sample is drawn as a break, not as a zero.

Physiological files arrive full of gaps: one of the recordings in our own test set is 99.7 per cent absent, a trigger with a hundred and thirty real marks in forty-four thousand rows. A dropped sample is not a measurement of zero, and drawing it as one puts a spike down to the floor of a trace that sits nowhere near zero, which reads as an artefact in the data. The viewer breaks the line instead. The filters still see a continuous signal, because they have to.

A trigger recording that is almost entirely gaps, drawn as short marks separated by empty space rather than as a line at zero
What a mostly-absent recording looks like. An external trigger with a hundred and thirty real marks in forty-four thousand rows, shown whole. The empty stretches are empty because nothing was recorded there, which is the honest picture; drawn as zeros it would be a solid line with spikes, and it would look like data.
Step 7

Look at the images.

Open the dataset in the Editor and click a converted image. It opens in the viewer: one plane at a time, or Multi-Planar for sagittal, coronal and axial sharing a single crosshair, where clicking or dragging in any plane moves the other two.

The image viewer. Scroll to move through slices, set brightness and contrast for a volume that looks flat, and set the crosshair colour and thickness to something you can see against your data.
A JSON sidecar as a schema-aware form with values read from the DICOM header
What the DICOM header already answered. Repetition and echo time, flip angle, field strength, coil combination, institution: all read out of the source and written for you. The handful of TODO values are the fields no file could supply, and they are left visible on purpose.

Check a functional run's time course

A 4-D file gains a Graph tab: click a voxel and see its value across every volume. Drift, spikes and motion are visible here in seconds, and it is worth doing on at least one run per session before you consider the conversion finished.

The 4-D graph. The signal at the crosshair voxel across every volume, beside the slices, without leaving the application.

And in three dimensions

The 3D tab renders the volume on the graphics card: rotate it, zoom, and cut through it with an oblique clipping plane whose cut face shows the slice underneath. It is the quickest way to spot a coverage problem or a wrap that is hard to see slice by slice. On a machine without suitable graphics support the tab is simply absent and the rest of the viewer works normally.

The 3-D renderer. Materials and lighting presets, an orientation cube, an oblique clip plane, and a switch between radiological and neurological convention that applies to the 2-D views as well.
Step 8

Validate.

Press Validate dataset. On this dataset: 55 clean files, 7 warnings, no errors. The warnings are the two kinds you should expect from MRI. The dataset-level fields, licence and authors, which are about your study. And the task descriptions and instructions on the functional sidecars, which are about your experiment. Both are things only you can supply, and both are answered in one pass.

Each finding names the schema rule it came from, so a requirement can be checked against the standard rather than taken on trust, and the fix button takes you to the field that needs the answer.

The validation pane showing the dataset, folder and file scopes each with a count
Three scopes, one pane. What is wrong with the dataset, with this folder, and with this file, each with its own count. These are not the same three as the toolbar buttons: those choose how much gets checked, these say what each finding is about. The tree on the left carries the per-file half of the same information, so the problem files, and the worst of them, are findable without opening every one.
A validation error on a scans table, with the schema rule printed underneath
An error inside a file. A timestamp in a scans table that is not a valid datetime. No filename check could see this, and no eye would catch it across seventy files. The rule path underneath, rules.tabular_data.modality_agnostic.Scans, is where the requirement comes from, and Highlight in editor tints the offending cell in the table.
Warnings on dataset_description.json, each with its schema rule and a fix button
What a warning looks like. Recommended fields nobody has answered yet, here on the dataset description, each with the standard's description of what it is for, an example of the value expected, and the rule it comes from. Fix opens the field. The amber pill beside the filename counts how many the file holds.
A warning that a task run has no events table beside it
A file that should be there and is not. The recording converted correctly and nothing about it is malformed. What is missing is the table describing what happened during it, which is the file most analyses will want. A warning rather than an error, because the standard recommends it rather than requires it, and the message says how to settle it either way: supply the table, or say in the task name that the run was rest.
Turn on Deep checks at least once.

The fast pass reads names and metadata. Deep checks additionally opens the NIfTI image headers, which is what catches a truncated or corrupt image. A structurally perfect dataset containing one unreadable volume passes every other kind of check.

A valid dataset is not a shareable one.

Nothing the validator checks has anything to say about the face in sub-001_T1w.nii.gz. If this dataset is going to anyone outside the people who collected it, deface it, and look at the result rather than trusting it: too little removed leaves the participant identifiable, too much takes the front of the brain with it. Defacing and skull stripping →

What you should see

Real numbers from this dataset.

Scan 33 / 12 rows / skipped

21 keepers across both subjects, 12 skipped automatically: localisers, report objects and calibrations. With probe convert on, about 25 seconds.

Convert 24 + 3 images + physio tables

24 NIfTI files across anat, func, dwi and fmap, three physiological tables from the embedded logs, three gradient table pairs, and 29 sidecars. No failures.

Validate 16 / 200 / 0 ok / warn / err

No errors. Counts are per finding, not per file, so 160 of those warnings are one TODO placeholder each, in a recommended field nobody has answered yet. Another 12 are functional runs with no events table beside them.

The same thing, scripted

From the command line.

# 1. Create the dataset and its project.
bidsmgr-create ~/bids/neuroimaging_unit --name "Neuroimaging unit"

# 2. Scan. 33 rows, 12 skipped. --probe-convert costs ~25 s and
#    saves a great deal of metadata typing later.
bidsmgr-scan ~/raw/neuroimaging_unit \
    --project ~/bids/neuroimaging_unit --probe-convert

# 3. Convert. 21 images plus one physiological table.
bidsmgr-convert --project ~/bids/neuroimaging_unit

# 4. Enrich: fieldmap IntendedFor, scans tables, participants.
bidsmgr-metadata --project ~/bids/neuroimaging_unit --fill-todos

# 5. Validate. 55 ok / 7 warn / 0 err.
bidsmgr-validate ~/bids/neuroimaging_unit

Site-specific classifier rules written in the interface are read here too, with --rules-file. Every flag is documented in the CLI reference.

Troubleshooting

What usually goes wrong with MRI.

What you seeWhat it means
Three subjects where you expected two The patient identifiers did not match, usually because the site anonymises per export. Edit the id cells; the grouping is a starting point, not a verdict.
One subject where you expected many The opposite: a reused or blank identifier. The same fix.
A series classified as something it is not Read the conf column. A low value means a name-based guess. Correct the row, and if it will recur, add a scan rule so the next scan gets it right.
A teal row with a crossed circle An object with no image data inside: TENSOR, a scanner report. Nothing can convert it. It is excluded on purpose, with the reason shown. Spectroscopy is not in this group: it converts to NIfTI-MRS and lands in mrs/.
Far more runs than you acquired Residual single-volume copies from dcm2niix were kept. They are dropped by default; check the setting.
A fieldmap with no IntendedFor Nothing matched it: usually the parameters differ from every candidate run, or the fieldmap sits in a different session. Set it by hand in the Editor.
Two task names that differ only in case Two tasks, as far as BIDS and every downstream pipeline are concerned. Fix it before converting; bulk edit does it in one move.

Next

The advanced MRI tutorial takes a richer protocol: several anatomical contrasts, scanner-derived diffusion maps, and a physiological log that fails at source. The multimodal tutorial shows MRI converting alongside PET, EEG and MEG in one pass.

Defacing and skull stripping is the page to read before this dataset leaves your group: the three ways to remove a face, how to check the result in the linked before-and-after viewer, and how to put the face back if the defacer took too much.

The full GUI walkthrough → CLI reference