§9.16–9.17Eight Million Higgs Bosons — Mass and Width

Part III Bettini pp. 416–421 · ~14 min read

  • signal strength
  • vector-boson fusion
  • narrow-width approximation

A width of 4 MeV is measured by an instrument whose resolution is 2 GeV — five hundred times coarser — because the number is pulled out of something other than the peak’s shape.

🎯 Why this matters

The move generalises: when a quantity sits below your resolution, stop measuring it directly and find an observable whose dependence on it is different. The instrument’s limit then stops being the measurement’s limit.

§9.15 found a boson on about a dozen four-lepton events. Run 2 delivered 150 fb⁻¹ at 13 TeV — five times the discovery dataset, and roughly eight million Higgs bosons. This section is what that buys: the ability to separate the production modes and decay channels from each other, and then to measure the mass to a part in a thousand and the width to a factor of two.

§9.16 What eight million Higgs bosons let you do

With the discovery sample you could see a peak. With Run 2 you can ask how each Higgs was made and how it decayed, separately — because the production modes have different kinematic signatures and can be told apart event by event.

Fig. 9.57 — the leading production processes, in order of importance

timeggHqqq'q'HqW, ZHggt t̄HggH — 87 %VBFVHttH, bbH

Click a vertex or an internal line.

Bettini Fig. 9.57. The measured quantity is a cross-section — the square of the SUM of these — but the experiments succeed in separating them, because each leaves a different set of extra objects in the event. That separation is what makes a coupling measurement possible at all.

💡 What this really says — why separating the production modes is the whole point

It would be easy to read Fig. 9.57 as a catalogue. It is not — it is the reason Run 2 could do something Run 1 could not.

Here is the problem. What you observe is a rate, and a rate is σproduction×BRdecay\sigma_{\text{production}} \times \text{BR}_{\text{decay}} — one number containing two couplings multiplied together. From one channel you cannot tell a 20 % excess in production from a 20 % excess in decay.

Measure the same decay in several different production modes, though, and the decay coupling is common to all of them while the production couplings differ. The system of equations becomes solvable. That is why the LHC quotes a grid of measurements — production mode × decay channel — rather than a list.

An engineer will recognise it as separating a product into its factors by varying one at a time, which is what a designed experiment does and what a single measurement never can. It is also why each production mode’s distinctive tag matters so much:

  • VBF leaves two forward jets, from the spectator quarks;
  • VH leaves a lepton from the vector boson — and that tag is what finally made HbbˉH\to b\bar b observable, after §9.13 showed it was hopeless inclusively;
  • ttH leaves a top pair, and is the only mode that reaches the top Yukawa directly rather than through the loop of §9.15.

Everything the experiments say about couplings rests on those tags.

The agreement is quantified by the signal strength μ\mu — the observed yield divided by the Standard Model prediction for it, defined so that μ=1\mu = 1 is agreement. One is quoted per production mode and per decay channel.

§9.17 The mass, from the two channels that give a peak

Bettini pp. 418–421. Note that the two channels' signal strengths agree with each other and with 1, which is the actual result.
channelwhat is measuredATLASCMS
HγγH\to\gamma\gammafiducial σ×\sigma\timesBR65.2±4.5stat±5.6syst±0.3theo65.2 \pm 4.5_{\text{stat}} \pm 5.6_{\text{syst}} \pm 0.3_{\text{theo}} fb, against a prediction of 63.6±3.363.6\pm3.3 fb
HγγH\to\gamma\gammasignal strength μ\mu1.02±0.141.02 \pm 0.141.12±0.091.12 \pm 0.09
HZZ4H\to ZZ^*\to4\ellsignal strength μ\mu1.01±0.111.01 \pm 0.110.940.11+0.120.94^{+0.12}_{-0.11}
HZZ4H\to ZZ^*\to4\ellthe mass124.99±0.17stat±0.03syst124.99 \pm 0.17_{\text{stat}} \pm 0.03_{\text{syst}} GeV125.38±0.14\mathbf{125.38 \pm 0.14} GeV

The four-lepton spectrum is worth drawing, because it contains something the di-photon spectrum does not — two peaks:

Fig. 9.60 — the CMS four-lepton spectrum, 137 fb⁻¹ at 13 TeV

4ℓ

80100120140160050100m(4ℓ) (GeV)events / 2 GeV

Two peaks. The one at 91 GeV is the RARE decay Z → 4ℓ — a single Z producing four leptons through an internal conversion — and it is not background to be subtracted but a standard candle: its position and width calibrate the lepton energy scale and the mass resolution on the very same final state as the signal. The peak at 125 GeV is the Higgs.

⚙️ Engineer’s bridge — the other peak is the calibration, and that is not a coincidence

The Z4Z\to4\ell peak is the best thing in this figure, and the book mentions it only in passing as “the rare decay of the Z in four leptons”.

Think about what it gives you. It is the same final state as the signal — four charged leptons — reconstructed by the same algorithms with the same efficiencies and the same resolution function, at a mass that is known to 23 ppm from LEP (§9.9). So:

  • if your reconstructed ZZ peak sits at the wrong mass, your lepton energy scale is wrong, and by exactly that fraction;
  • if its width is wider than expected, your resolution model is wrong, and by exactly that factor.

Both corrections then transfer directly to the Higgs peak 34 GeV away.

This is the §9.11 trick again — the Tevatron calibrated its lepton scale on ZZ\to\ell\ell and its jet scale on the hadronic WW — and it is the same instinct an engineer applies with a known reference in the same channel: a calibration tone inside the measurement band, a pilot carrier, a resistor of known value in the arm you are not measuring. The rule is that a reference which shares the signal’s path removes everything the path does to it.

It is also why the ATLAS and CMS mass measurements can be quoted with a systematic error of 0.03 GeV against a statistical error of 0.17: almost everything that could go wrong systematically has been measured away on a peak sitting in the same plot.

Where it breaks: an in-situ standard calibrates over the range it spans. The Z sits at 91 GeV and the diphoton peak at 125, so using one for the other means extrapolating the energy scale by nearly 40 % — and non-linearity is exactly the error a single-point standard cannot see, because a scale error and a linearity error are degenerate until you have two points. It is also channel- specific: ZeeZ \to ee calibrates electrons, and photons differ in how they shower and in how much material they have converted in. The calibration peak removes the systematics it shares with the signal and is silent about the rest.

Erratum — the four-lepton branching ratio is quoted 15× too large

The book writes: “about 2.7 % HZZH\to ZZ^* times 6.7 % for the decay of the ZZ into e+ee^+e^- or μ+μ\mu^+\mu^- namely 0.18 %.”

The arithmetic 2.7%×6.7%=0.18%2.7\,\% \times 6.7\,\% = 0.18\,\% is right, but both ZZs must decay to leptons, so the factor of 6.7 % belongs twice:

BR=0.0262×(0.0673)2=1.19×104=0.012%\text{BR} = 0.0262 \times (0.0673)^2 = 1.19\times10^{-4} = \mathbf{0.012\,\%}

The book’s own numbers confirm it. §9.15 quotes σ×BR=2.8\sigma\times\text{BR} = 2.8 fb for this channel at 8 TeV against a total σH=22.3\sigma_H = 22.3 pb, and 2.8/22300=1.26×1042.8/22300 = 1.26\times10^{-4} — agreeing with 0.012%0.012\,\%, not with 0.18%0.18\,\%.

It matters, because a factor of 15 in this branching ratio is the difference between expecting hundreds of four-lepton events and expecting about a dozen — and about a dozen is what §9.15 shows the discovery actually rested on. Confirmed on the render of PDF p. 437.

The width: 4 MeV, measured with 2 GeV resolution

Everything above measured a position. The width is a different problem entirely, and the book’s treatment of it — through the narrow-width approximation — is the cleverest thing in the chapter.

why the width cannot be measured, and then how it is

G, M = 4.14e-3, 125.38          # GeV
hbar = 6.582119569e-25          # GeV s
print("the Standard Model prediction")
print(f"  Gamma_H  = {G*1e3:.2f} +- 0.02 MeV")
print(f"  tau_H    = hbar / Gamma = {hbar/G:.2e} s")
print("\nwhy you cannot see it directly:")
print(f"  Gamma / M = {G*1e3:.2f} MeV / {M*1e3:.0f} MeV = {G/M:.1e}")
print(f"  so you would need a relative energy resolution better than 3e-5")
print(f"\n  what the experiments actually have: 1-2 GeV on 125 GeV = {1.5/M:.1e}")
print(f"  that is {(1.5/M)/(G/M):.0f}x too coarse.  the observed peak width is ENTIRELY")
print( "  instrumental -- the physics width contributes nothing to it.")

print("\nso the width is measured a completely different way, off shell:")
print( "  on-shell  yield  ~  g_p^2 g_d^2 / Gamma_H")
print( "  off-shell yield  ~  g_p^2 g_d^2")
print("\n  the 1/Gamma comes from the resonance integral, and the off-shell")
print( "  process does not resonate, so it does not have one.  the couplings")
print( "  are the SAME in both, so their ratio is Gamma_H and nothing else.")
print("\n  CMS result: Gamma_H = 3.2 (+2.4 -1.7) MeV")
print(f"  the SM says {G*1e3:.2f}.  agreement, at a factor of two.")

print("\nthe resonance integral that supplies the 1/Gamma:")
print( "  int d(s-hat) / [(s-hat - M^2)^2 + M^2 Gamma^2] = pi / (M Gamma)")
print( "  -> the Breit-Wigner acts like  (pi/(M Gamma)) x delta(s-hat - M^2)")
print( "  squeeze a curve of fixed AREA and its height rises as 1/width.")
prints
the Standard Model prediction
Gamma_H  = 4.14 +- 0.02 MeV
tau_H    = hbar / Gamma = 1.59e-22 s

why you cannot see it directly:
Gamma / M = 4.14 MeV / 125380 MeV = 3.3e-05
so you would need a relative energy resolution better than 3e-5

what the experiments actually have: 1-2 GeV on 125 GeV = 1.2e-02
that is 362x too coarse.  the observed peak width is ENTIRELY
instrumental -- the physics width contributes nothing to it.

so the width is measured a completely different way, off shell:
on-shell  yield  ~  g_p^2 g_d^2 / Gamma_H
off-shell yield  ~  g_p^2 g_d^2

the 1/Gamma comes from the resonance integral, and the off-shell
process does not resonate, so it does not have one.  the couplings
are the SAME in both, so their ratio is Gamma_H and nothing else.

CMS result: Gamma_H = 3.2 (+2.4 -1.7) MeV
the SM says 4.14.  agreement, at a factor of two.

the resonance integral that supplies the 1/Gamma:
int d(s-hat) / [(s-hat - M^2)^2 + M^2 Gamma^2] = pi / (M Gamma)
-> the Breit-Wigner acts like  (pi/(M Gamma)) x delta(s-hat - M^2)
squeeze a curve of fixed AREA and its height rises as 1/width.
1(s^MH2)2+MH2ΓH2    πMHΓHδ ⁣(s^MH2)\frac{1}{\left(\hat s - M_H^2\right)^2 + M_H^2\Gamma_H^2} \;\longrightarrow\; \htmlClass{t-c}{\frac{\pi}{M_H\Gamma_H}}\, \htmlClass{t-d}{\delta\!\left(\hat s - M_H^2\right)}
(9.128)

Bettini p. 421, the narrow-width approximation — and the whole width measurement lives in the coefficient, not in the delta function.

Every symbol, one at a time

Hover or tap a symbol above — it lights up in the equation and its meaning, units and type appear here.

🪜 Measuring a 4 MeV width without resolving it — Eqs. (9.127)–(9.130)

Step 1 of 4the narrow-width approximation

ds^(s^MH2)2+MH2ΓH2=πMHΓH\int \frac{d\hat s}{\left(\hat s - M_H^2\right)^2 + M_H^2\Gamma_H^2} = \frac{\pi}{M_H\Gamma_H}

Why you may do this: Γ_H is so small that nothing else in the problem varies across the width, so the resonance can be replaced by a delta function — but a delta function with a coefficient, and the coefficient contains 1/Γ.

The substitution x = (ŝ − M²)/(M Γ) turns the integral into ∫dx/(1+x²) = π. Everything else is bookkeeping.

Bettini pp. 421–422. The whole argument is that one of the two processes carries a 1/Γ and the other does not.

💡 What this really says — measuring a quantity through the one thing that depends on it

Step back from the algebra, because the strategy generalises far beyond this measurement.

You want ΓH\Gamma_H. It appears in exactly one observable feature — the height of the resonance relative to what the couplings alone would give. But you cannot isolate that, because the observed height also depends on the couplings, which you do not independently know.

The fix is to find a second measurement in which the couplings appear the same way and the width does not. Then the ratio has the couplings cancel and the width survive. The off-shell region is exactly that: same vertices, same particles, no resonance.

An engineer will recognise the structure as a two-measurement solve for a nuisance-coupled parameter — and specifically as the trick behind ratiometric measurement. You cannot measure a resistance with an unknown excitation current; measure two resistors with the same current and the ratio is exact. Here the “unknown current” is gp2gd2g_p^2g_d^2 and the two “resistors” are the on-shell and off-shell regimes.

Two things make it work and both are worth noticing. First, the couplings must genuinely be the same — which they are, because it is literally the same vertex at a different s^\hat s. Second, the off-shell rate must be observable at all, which it barely is: it needed the whole of Run 2. The result, 3.21.7+2.43.2^{+2.4}_{-1.7} MeV against a predicted 4.14, is a factor-of-two measurement of a quantity 400 times finer than the apparatus can resolve.

🔑 If you remember only three things

  • Run 1 found it and Run 2 asked what it is. The difference is not precision so much as being able to put separate questions to separate production modes.

  • Eight million is a number about questions, not about error bars. One million would have measured the same mass and answered fewer things.

  • A known peak beside the unknown one removes the energy scale. The calibration travels in the same histogram as the measurement, which no separate run could achieve.

Where this goes next

The mass is known to 0.14 GeV and the width to a factor of two, and every signal strength measured so far sits on 1. What remains is the part that distinguishes this scalar from any other:

§9.18 establishes JP=0+J^P = 0^+ from the angle between the two ZZ decay planes — the argument §3.5 used on the π0\pi^0, with massive vector bosons in place of photons. §9.19 then measures the couplings against mass and looks for the two power laws §9.12 predicts: linear for fermions, quadratic for bosons, both through the same υ\upsilon.

Check yourself — eight million Higgs bosons

0/6 answered · 0 correct

  1. 1.Why does separating the production modes matter so much, rather than just counting Higgs bosons?

  2. 2.The four-lepton spectrum has a peak at 91 GeV as well as one at 125. What is the first one, and why is it valuable?

  3. 3.The book says BR(H → ZZ* → 4ℓ) is 'about 2.7 % times 6.7 % namely 0.18 %'. What is wrong?

  4. 4.Γ_H = 4.1 MeV and the mass resolution is 1–2 GeV. How is the width measured at all?

  5. 5.In the narrow-width approximation the Breit–Wigner becomes π/(M Γ) × δ(ŝ − M²). Where does the physics live?

  6. 6.Every signal strength quoted in this section sits within ~10 % of 1. Does that establish that the particle is the Standard Model Higgs?

Study aid derived from A. Bettini, Introduction to Elementary Particle Physics, 3rd ed., Cambridge University Press 2024 — published Open Access under CC-BY-NC 4.0, DOI 10.1017/9781009440745. Not the book: an independently written interactive companion, figures redrawn.