Finding the first quasars in Euclid

A pipeline and a web console that search the Euclid survey for quasars at redshift 7 and beyond.

The problem

Euclid is imaging about 15,000 square degrees in one optical band and three near-infrared bands. Somewhere in that footprint are the most distant quasars anyone will find this decade, at redshifts where the Universe was under 800 million years old. At the survey's depth there is roughly one of them per thousand square degrees.

Brown dwarfs in our own galaxy have almost the same colours in those four bands and outnumber the quasars by hundreds to one. Every candidate list is therefore a contamination problem before it is a discovery problem. It is the same statistics as finding extremely metal-poor stars in a spectroscopic survey, which was my PhD: a rare class, a common look-alike, and a decision that has to be made from a handful of noisy numbers.

What I built

An end-to-end search, from raw survey tiles to a vetted candidate list, and a console for looking at every candidate properly.

  • The crawl. Every released Euclid tile, 8,528 of them, acquired and processed by parallel workers at around 250 tiles a day. The system runs unattended and recovers by itself: watchdog processes restart the workers and the console, and the crawl resumes from where it stopped after an unplanned outage.
  • Bayesian scoring, not colour cuts. For every source, the photometry is compared against quasar templates across redshift and against a library of M, L and T dwarf and galaxy templates. The output is a posterior probability that the source is a high-redshift quasar, with a photometric redshift, and the contaminants are modelled explicitly rather than cut around. External surveys are folded in where they cover the field.
  • A gauntlet of vetoes. A z > 7 quasar must be invisible in the optical band. The optical veto uses aperture signal-to-noise on the median-subtracted stack, calibrated against 9,589 blank-sky positions: it was the only test that discriminated, since pixel-peak and streak tests fired on empty sky just as often as on real sources. Further flags mark faint red objects, sources near bright stars, and tile edges.
  • Validation against what is known. Published high-redshift quasars, their known contaminants and a lower-redshift quasar catalogue are cross-matched every six hours. Known quasars in the footprint are recovered blind at probabilities near one. One known quasar is scored low because a T-dwarf template fits its colours better; that is a genuine degeneracy in four-band photometry, documented rather than patched.
  • Selection function and purity. Injection-recovery grids over luminosity and redshift with a hundred realisations per cell give completeness as a function of magnitude. Purity comes from 24 dwarf spectral types and two galaxy templates injected at fifty realisations each. Predicted yields land slightly above the published Euclid forecasts in every magnitude bin.

The console

A web app for vetting candidates, because a probability is not a decision. Each candidate gets a page with:

  • cutouts from every survey that covers it, ordered by wavelength: Pan-STARRS or Legacy Survey in the optical, Euclid itself, unWISE in the mid-infrared, GALEX in the ultraviolet, VLASS and LoTSS in the radio, HSC where it exists;
  • the spectral energy distribution against the best quasar and best dwarf templates;
  • an automatic check of the Hubble archive for existing observations;
  • a vetting column, good, bad or unsure, whose labels are accumulating into a training set for a classifier;
  • export as CSV with a magnitude cut, and a one-page PDF sheet per candidate for sharing with collaborators.

The views are served from DuckDB with a caching layer, scoring is incremental so only new sources are recomputed, and the expensive views are kept warm between requests.

What is next

Moving from the mosaics to Euclid's per-observation stacked frames, which keep detail the mosaics lose, and training a classifier on the human vetting labels so the console can rank what a person should look at first.