One demand — that the equation be first order in time, as every other equation of motion is — forces coefficients that cannot be numbers. The rest of the page works out what they must be.
🎯 Why this matters
Everything after this page is written in the notation the page forces: fields with four components and matrices between them. The algebra settled here is the vocabulary the rest of the book uses without ever re-deriving it.This is the most theoretical section in the first half of the book, and it is also the one where a computer engineer has the biggest advantage: from here to the end of the section, everything is linear algebra. Anticommutators, a Clifford algebra, a change of basis, a complete basis for a matrix space, an eigenvalue projection. The physics is in what the algebra is forced to do.
The result is worth stating before the derivation, because it is one of the most remarkable in the subject. Dirac wanted an equation that was first order in time. That single requirement forces the coefficients to anticommute, which forces them to be matrices, which forces the wave function to have four components — two of which turn out to be the two spin states, and two of which turn out to be the antiparticle. Nobody put spin in. Nobody put antimatter in. They fall out of demanding relativistic covariance and a first-order equation.
Why the obvious relativistic equation fails
📐 Physics you need first — what a wave equation is doing here
Quantum mechanics represents a particle by a complex function whose squared modulus is a probability density: is the chance of finding the particle in . Its evolution comes from taking the classical energy–momentum relation and replacing
(natural units, ). Do that to the non-relativistic and you get the Schrödinger equation. Two properties matter for what follows:
- the equation is first order in , so specifying at one instant determines it for ever after;
- is conserved and non-negative, which is what lets you call it a probability at all.
The relativistic relation is , and it is quadratic. Substituting gives an equation with a second time derivative — the Klein–Gordon equation Klein–Gordon equation the relativistic wave equation obtained from E² = p² + m²; Dirac rejected it because a second-order time derivative lets the probability density go negative. defined in §2.5 — open in glossary — and both properties above break.
The Klein–Gordon equation. It is perfectly correct — it is the right equation for a spin-0 particle, and Chapter 5 uses it — but it cannot describe an electron, and Dirac's reason for rejecting it was the probability interpretation.
Every symbol, one at a time
Hover or tap a symbol above — it lights up in the equation and its meaning, units and type appear here.
⚙️ Engineer’s bridge — Dirac wanted a state-space form
Read the two equations as dynamical systems and the objection becomes familiar.
First order in time is : the state at one instant determines the whole trajectory, and “the state” is a well-defined object you can carry around. That is what Schrödinger’s equation is, with .
Second order in time is . You can always convert it to first order — that is what a state-space realisation does — but only by doubling the state vector to . Nothing is wrong with that in control theory. In quantum mechanics it is fatal, because the quantity you wanted to interpret as a probability is built from alone, and once is independently specifiable you can choose initial conditions that make it negative.
So Dirac’s demand is: give me a genuine first-order system, of the form
with containing to the first power. The rest of this section is what it costs to get one.
Where the analogy breaks: converting Klein–Gordon to first order by doubling the state is a perfectly good move mathematically, and modern field theory does exactly that. The objection is specifically to the single-particle probability interpretation, which quantum field theory later abandons anyway. Dirac was solving the right problem for the wrong reason, and got the right answer.
Where it breaks: reducing an equation to first order in time is a representational move in control theory — you trade one second-order state for two first-order ones and nothing about the system changes. Here it is not representational. Making the Dirac equation first order forces anticommuting coefficients, hence 4 × 4 matrices, hence a four-component object with negative- energy solutions that were not in the problem you started with — and those solutions turned out to be antimatter. A state-space rewrite that predicts a new particle is not a change of variables, and Dirac’s motivation for wanting it (the probability interpretation) was abandoned by field theory anyway.
The ansatz, and what it forces
🪜 From “first order in time” to “the coefficients must be matrices”
Step 1 of 7 — Write the most general first-order Hamiltonian(2.28)
Why you may do this: Linear in momentum, linear in mass, with four unknown coefficients α₁, α₂, α₃ and β. Nothing yet says what kind of objects they are — that is what the next steps discover.
Repeated Latin indices are summed 1…3 throughout. This is just "the most general thing that is first order".
Bettini Eqs. (2.28)–(2.34). Every step is algebra; the physics is entirely in the demand of step 1.
⚙️ Engineer’s bridge — this is a Clifford algebra, and the size is a dimension count
Strip the physics away and the problem is: find matrices with and for . That is the definition of a Clifford algebra, and the smallest matrix size that admits such generators is fixed, not negotiable:
| generators needed | smallest matrices | what it describes |
|---|---|---|
| 2 | 2 × 2 | — |
| 3 | 2 × 2 | the Pauli matrices — spin in three dimensions |
| 4 | 4 × 4 | the Dirac matrices — spin in 3+1 space-time |
| 5, 6 | 8 × 8 | (five dimensions, and beyond) |
The pattern is . Three generators fit in 2×2 with nothing to spare; the fourth doubles the size. So the four-component spinor is not a modelling choice — it is the smallest representation the algebra admits, and the number four is dictated by the number of space-time dimensions.
The dimension counting is exactly the argument you would use for any generator set: how many independent objects satisfying these relations can live in a space of this size? Here the answer decides how many components an electron has.
Where it breaks: dimension counting fixes the minimum size of the matrices and does not tell you which representation nature uses. The same Clifford algebra admits the Weyl form (two two-component spinors) and the Majorana form (a four-component real spinor describing a particle identical to its antiparticle) — §2.9 is about exactly that ambiguity, and whether the neutrino takes the Majorana option is still unknown. The algebra constrains the container; it does not choose the contents.
🔢 The Dirac algebra, as matrices
Dirac: The book’s choice, Eq. (2.36). γ⁰ is diagonal, which makes the non-relativistic limit transparent: the upper two components are the particle and the lower two the antiparticle, and at low energy they decouple.
| 1 | 0 | 0 | 0 |
| 0 | 1 | 0 | 0 |
| 0 | 0 | -1 | 0 |
| 0 | 0 | 0 | -1 |
| 0 | 0 | 0 | 1 |
| 0 | 0 | 1 | 0 |
| 0 | -1 | 0 | 0 |
| -1 | 0 | 0 | 0 |
| 0 | 0 | 0 | −i |
| 0 | 0 | i | 0 |
| 0 | i | 0 | 0 |
| −i | 0 | 0 | 0 |
| 0 | 0 | 1 | 0 |
| 0 | 0 | 0 | -1 |
| -1 | 0 | 0 | 0 |
| 0 | 1 | 0 | 0 |
| 0 | 0 | 1 | 0 |
| 0 | 0 | 0 | 1 |
| 1 | 0 | 0 | 0 |
| 0 | 1 | 0 | 0 |
Check the defining relation yourself
| 0 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 |
2 g01 · 𝟙 = 2 × 0 × 𝟙
✓ they agree, in every representation
Try all sixteen pairs, then switch representation and try again. The matrices change completely; this relation does not. That is what it means to say the choice of representation carries no physics.
Why exactly five bilinear covariants — a dimension count
| Covariant | Transforms as | Independent components |
|---|---|---|
| ψ̄ψ | scalar | 1 |
| ψ̄γ⁵ψ | pseudoscalar | 1 |
| ψ̄γ^μψ | vector | 4 |
| ψ̄γ^μγ⁵ψ | axial vector | 4 |
| ψ̄σ^μνψ | tensor | 6 |
| total | 16 | |
Sixteen is 4² — the dimension of the space of all 4×4 matrices. So the five covariants are not a list somebody collected: they are a complete basis, sorted into the pieces that transform into themselves under Lorentz transformations. Any 4×4 matrix you could put between ψ̄ and ψ is a combination of these, which is why the list stops at five. Nature then uses only two of them, V and A — and that is a physical fact, not a mathematical one (§2.8, Ch. 7).
The equation
Substituting (2.26) into the ansatz gives (2.35). Define the gamma matrices gamma matrices the four 4×4 matrices obeying {γ^μ, γ^ν} = 2g^μν; the Dirac representation is one choice among many, and any set satisfying the anticommutation relation will do. defined in §2.5 — open in glossary and (2.36), multiply through by from the left, and every index falls into place:
The Dirac equation. It is Lorentz covariant, first order in time, and its conserved density is positive — which is everything Dirac asked for.
Every symbol, one at a time
Hover or tap a symbol above — it lights up in the equation and its meaning, units and type appear here.
Solid arrows are forced steps; dashed arrows are the physics that arrives uninvited. Three assumptions go in on the left and four results come out on the right — and none of the three mentions spin, magnetic moments or antiparticles.
💡 What this really says — what you got for free
Count what was assumed and what emerged.
Assumed: relativity (), the quantum substitution rule, and first order in . Three things, none of them about spin or antimatter.
Emerged:
- Four components. The smallest matrices that satisfy the algebra are 4 × 4, so the wave function has four entries.
- Spin. The α matrices are built out of the Pauli matrices, and the Pauli matrices are the mathematical representation of spin ½. So the equation describes a particle with intrinsic angular momentum, and it could not have described anything else. As the book puts it, spin is fundamentally required by Lorentz invariance — not an extra property bolted on.
- The magnetic moment, with the right size. The equation predicts a gyromagnetic ratio of exactly (2.39), where classical reasoning would give . No parameter was fitted.
- Antimatter. Two of the four components describe a state of the opposite charge. Anderson found it in 1932 (§2.6–2.7), four years after Dirac wrote the equation.
That is a large amount of physics obtained from a structural demand, and it is the reason this equation is treated with the reverence it is.
🔢 Worked example — g = 2, and the number it predicts
The magnetic moment of a spin-½ particle is written
with the Bohr magneton. Classically, a charge distribution rotating with angular momentum gives . The Dirac equation gives , with nothing adjustable.
Put in numbers. An electron in a 1 T field has its two spin states split by
which is 28.0 GHz in frequency units — a microwave transition, and exactly what electron spin resonance measures.
The measured value is not quite 2: it is . That part per thousand is a quantum-electrodynamic correction, calculated and measured to twelve digits, and it is the most precisely verified prediction in physics (Ch. 5). The Dirac equation gets the leading term exactly right without trying.
The rest of the machinery
Everything below is bookkeeping you will need in Chapters 5 and 7, and it is all linear algebra.
| Object | Eq.↕ | What it is |
|---|---|---|
| 2.42 | The bispinor: four complex components as two stacked two-component spinors. φ is the particle and χ the antiparticle; within each, the two entries are s_z = +½ and −½. | |
| 2.43 | ||
| 2.45 | ||
| 2.48 | A fifth matrix, built from the other four. (γ⁵)² = 1 and γ⁵† = γ⁵, so its eigenvalues are ±1 — and those eigenvalues are the CHIRALITY of §2.8. | |
| 2.51 | The Dirac Lagrangian density. Feed it to the Euler–Lagrange equation and the Dirac equation comes back out; vary with respect to ψ instead of ψ̄ and you get the adjoint equation. |
The Lagrangian is the form everything later is written in, because <strong>interactions are added by adding terms to it</strong>. Ch. 5 adds one term and gets all of QED; Ch. 9 adds a few more and gets the Standard Model. This one line is the template.
The bispinor of Eq. (2.42) carries two independent binary labels — particle versus antiparticle, and spin up versus down — but chirality is a third split that does not line up with either. Switch the widget above to the chiral basis and the right-hand decomposition becomes the diagonal one instead.
📐 Physics you need first — what a Lagrangian density is for
In mechanics, instead of writing the equation of motion you can write a single scalar and derive the equation from it, by demanding that the action be stationary. The recipe is Euler–Lagrange:
Field theory does the same with the coordinate replaced by a field and replaced by (2.52). Treat and as independent, and each gives one of the two equations.
Why bother, when you already have the equation? Three reasons, all of which matter later:
- Lorentz invariance is manifest. is a scalar; if you write down only scalars, every equation you derive is automatically covariant.
- Symmetries become visible. Noether’s theorem reads a conservation law straight off an invariance of — this is where charge conservation comes from in Ch. 5.
- Interactions are additive. You do not modify the equation; you add a term to and turn the handle.
It is, in short, a specification you can compose, rather than a solution you have to re-derive.
Reproduce it
import numpy as np
I2 = np.eye(2, dtype=complex); Z = np.zeros((2, 2), dtype=complex)
s = [np.array([[0,1],[1,0]], dtype=complex),
np.array([[0,-1j],[1j,0]], dtype=complex),
np.array([[1,0],[0,-1]], dtype=complex)]
blk = lambda a,b,c,d: np.block([[a,b],[c,d]])
g = np.diag([1,-1,-1,-1]).astype(complex)
beta = blk(I2, Z, Z, -I2) # Eq. (2.34)
alpha = [blk(Z, x, x, Z) for x in s]
gD = [beta] + [beta @ a for a in alpha] # Eq. (2.36)
print("Dirac representation (2.34), (2.36): gamma^0 =")
for row in gD[0].real.astype(int):
print(" " + " ".join(f"{x:2d}" for x in row))
print(f"conditions on the ansatz: beta^2 = 1 : {np.allclose(beta@beta, np.eye(4))}")
print(f" {{alpha_k, beta}} = 0 : "
f"{all(np.allclose(a@beta+beta@a, 0) for a in alpha)}")
print(f" {{alpha_k, alpha_j}} = 2 delta_kj : "
f"{all(np.allclose(alpha[i]@alpha[j]+alpha[j]@alpha[i], 2*(i==j)*np.eye(4)) for i in range(3) for j in range(3))}")
gW = [blk(Z, I2, I2, Z)] + [blk(Z, x, -x, Z) for x in s] # chiral
kr = np.kron
gM = [kr(s[1],s[0]), 1j*kr(s[0],I2), 1j*kr(s[2],I2), 1j*kr(s[1],s[1])] # Eq. (2.67)
print("{gamma^mu, gamma^nu} = 2 g^munu holds in every representation:")
for name, G in (("Dirac", gD), ("chiral", gW), ("Majorana", gM)):
ok = all(np.allclose(G[m]@G[n] + G[n]@G[m], 2*g[m,n]*np.eye(4))
for m in range(4) for n in range(4))
print(f" {name:8s} {ok} (purely imaginary: {all(np.allclose(x.real, 0) for x in G)})")
g5 = 1j*gD[0]@gD[1]@gD[2]@gD[3] # Eq. (2.48)
print(f"gamma^5 = i g0 g1 g2 g3 : (g5)^2 = 1 {np.allclose(g5@g5, np.eye(4))}, "
f"hermitian {np.allclose(g5, g5.conj().T)}")
basis = ([np.eye(4), g5] + gD + [x @ g5 for x in gD]
+ [0.5j*(gD[m]@gD[n]-gD[n]@gD[m]) for m in range(4) for n in range(m+1,4)])
rank = np.linalg.matrix_rank(np.array([b.flatten() for b in basis]))
print(f"the 16 bilinear covariants span the 4x4 matrices: rank = {rank} of {len(basis)}")
g_meas = 1.00115965218
print(f"g = 2 predicted; measured g/2 = {g_meas}, so the QED correction is {g_meas-1:.1e}") Dirac representation (2.34), (2.36): gamma^0 =
1 0 0 0
0 1 0 0
0 0 -1 0
0 0 0 -1
conditions on the ansatz: beta^2 = 1 : True
{alpha_k, beta} = 0 : True
{alpha_k, alpha_j} = 2 delta_kj : True
{gamma^mu, gamma^nu} = 2 g^munu holds in every representation:
Dirac True (purely imaginary: False)
chiral True (purely imaginary: False)
Majorana True (purely imaginary: True)
gamma^5 = i g0 g1 g2 g3 : (g5)^2 = 1 True, hermitian True
the 16 bilinear covariants span the 4x4 matrices: rank = 16 of 16
g = 2 predicted; measured g/2 = 1.00115965218, so the QED correction is 1.2e-03 🔑 If you remember only three things
-
The matrix size was forced, not chosen. Four is the smallest dimension in which the algebra closes, so the four components of ψ are a result rather than a modelling decision.
-
A defect drove the whole construction. The equation was not built to describe spin; it was built to be first order, and what it turned out to contain arrived as a consequence.
-
g = 2 is the receipt. A derivation that only repaired a formal problem would have been a trick; predicting a measured number nobody asked for is what made it a theory.
Where this goes next
- §2.6–2.7 is the antiparticle this equation predicted, found in a cloud chamber four years later.
- §2.8–2.9 takes γ⁵ seriously: its eigenvalues are chirality, and the Majorana representation of the widget above is what a self-conjugate fermion needs.
- Ch. 5 adds one term to the Lagrangian above and gets QED, including the correction to g = 2.
- Ch. 7 is where “nature uses only V and A” out of the five bilinear covariants becomes the defining property of the weak interaction.
✅ Check yourself — what a first-order equation costs
0/6 answered · 0 correct
1.Dirac rejected the Klein–Gordon equation because it is second order in time. Read as a dynamical system, what is the objection?
2.Matching H² = p² + m² term by term gives β² = 1, {α_k, β} = 0 and {α_k, α_j} = 2δ_kj. Why do these force the coefficients to be matrices?
3.The Pauli matrices already give three mutually anticommuting 2 × 2 matrices squaring to 1. Why does a fourth need 4 × 4?
4.In the widget, switching between the Dirac, chiral and Majorana representations changes every matrix entry but never breaks {γ^μ, γ^ν} = 2g^μν. What does that mean?
5.There are exactly five bilinear covariants: scalar, pseudoscalar, vector, axial vector, tensor. Why five, and not four or six?
6.The Dirac equation predicts g = 2 for the electron, where classical reasoning gives g = 1. What is the status of that prediction?