MLB Team Defense Update (9/7/26)

Defensive statistics have always been difficult to interpret.

Batting statistics usually tell a fairly direct story. A hitter gets on base, hits for power, strikes out, walks, and produces runs. Pitching statistics are more complicated, but the broad questions are still familiar. Does the pitcher miss bats? Does he limit walks? Does he prevent runs?

Defense is different.

No single defensive statistic is universally accepted as definitive. Defensive Runs Saved, Outs Above Average, Fielding Run Value, FanGraphs Def, fielding percentage, and catcher throwing statistics all attempt to measure defense from slightly different directions. Sometimes they agree. Sometimes they disagree dramatically.

That makes team defense an ideal candidate for principal component analysis. Rather than deciding in advance which defensive statistic is “correct,” PCA lets the statistics reveal the dominant patterns in the data.

For this study, I used 2026 FanGraphs team defensive data through September 7. Six measures were included:

  • Defensive Runs Saved
  • Outs Above Average
  • Fielding Run Value
  • FanGraphs Def
  • Fielding percentage
  • Caught-stealing rate

I standardized each variable before performing the PCA. This is important because the statistics exist on very different numerical scales. Fielding percentage, for example, clusters around .980 to .990, while DRS can range from strongly negative to well over +100. Without standardization, a statistic’s scale could influence the PCA more than the information it contains.

The result was surprisingly clean. The first principal component explains 60.5 percent of all variation among MLB team defenses. Even more importantly, its meaning is easy to interpret.

The loadings on PC1 were:

\mathrm{Def} = 0.507 \mathrm{FRV} = 0.506 \mathrm{OAA} = 0.493 \mathrm{DRS} = 0.393 \mathrm{FP} = 0.297

Caught-stealing rate contributed almost nothing to the first component.

In practical terms, PC1 acts as an overall defensive quality axis. That interpretation is reinforced by the extremely strong relationship between PC1 and FanGraphs Def. The correlation is about 0.97. So although PCA was not told what constituted “good defense,” it independently produced a first component that behaves almost exactly like a composite defensive-quality measure.

And one team separates itself immediately.

Chicago is not merely first; the Cubs are in a different neighborhood. Their PC1 score is 5.94, compared with 2.81 for second-place Arizona. No other team approaches Chicago’s position on the primary defensive axis. That distance is important. Rankings can sometimes exaggerate small differences. A team ranked first may be only marginally better than the team ranked second.

That is not what is happening here. The PCA suggests an enormous separation between the Cubs and everyone else.

Chicago entered September 8 with 107 Defensive Runs Saved, 65 Outs Above Average, and 64 Fielding Run Value in the data used here. Those are not merely good numbers. They represent broad agreement among different defensive measurement systems that Chicago has been exceptional.

The top ten teams by PC1 were:

Rank Team PC1
1 Chicago Cubs 5.94
2 Arizona 2.81
3 St. Louis 2.01
4 Toronto 1.81
5 San Diego 1.68
6 Los Angeles Dodgers 1.62
7 Atlanta 1.61
8 Boston 1.33
9 Kansas City 1.18
10 Cleveland 0.84

Arizona emerges as the clear second-place team. The Diamondbacks’ defensive profile is particularly strong in the advanced range-based measures. They recorded 39 OAA and 31 FRV, producing a PC1 score substantially above most of the league.

St. Louis ranks third, followed by Toronto and San Diego. Cleveland comes in tenth. That is a respectable position, but the graph makes clear how far Chicago is from a good defensive team.

The opposite end of the PCA is equally interesting. Seattle ranks last with a PC1 score of -3.29. The Athletics are close behind at -3.02, followed by the Angels, Colorado, and Minnesota.

Rank Team PC1
30 Seattle -3.29
29 Athletics -3.02
28 Los Angeles Angels -2.33
27 Colorado -2.13
26 Minnesota -1.93
25 Pittsburgh -1.82
24 Cincinnati -1.46
23 San Francisco -1.41

Seattle’s placement is driven heavily by -51 OAA and -44 FRV. Those numbers suggest a club that has struggled considerably to convert balls in play into outs relative to what would be expected.

But the lower portion of the rankings also reveals why using multiple defensive measurements is useful. Cincinnati, for example, had +1 OAA but -45 DRS. San Francisco showed almost the opposite disagreement, with +22 DRS but -17 OAA.

Which statistic should we trust? That is precisely the wrong question. Different defensive systems use different models, assumptions, opportunities, positioning adjustments, and definitions of responsibility. Disagreement among them is therefore not necessarily evidence that one system has failed.

It can also reveal uncertainty.

PCA helps by asking a different question: across all of these measurements, what common defensive signal appears most consistently? For most teams, that signal is PC1.

The second principal component tells a completely different story. PC2 explains another 17.3 percent of the total variance, bringing the first two principal components to approximately 77.9 percent of all defensive variation in the six original variables.

But PC2 is not another general-defense measure. It reflects the running game almost entirely.

The loading for caught-stealing rate on PC2 is approximately: 0.961. That is extraordinarily large.

The other variables contribute comparatively little. This means that the vertical axis in Figure 1 can essentially be interpreted as running-game control, while the horizontal axis measures broader defensive quality.

That makes several teams particularly interesting. San Diego ranks fifth overall defensively, but the Padres also sit very high on PC2. Kansas City shows a similar pattern. Their location on the graph suggests a defensive identity that differs from teams such as Arizona or Atlanta.

Those clubs may all be good defensively, but not in exactly the same way. And that is one of PCA’s greatest strengths. A simple ranking compresses every team into one number. The PCA retains structure.

Teams can be similar in overall quality while achieving that quality through different defensive profiles. However, there is an important methodological limitation.

DRS, OAA, FRV, and FanGraphs Def are not completely independent measurements. Several are derived from overlapping types of defensive information. OAA and FRV in particular are closely related, while FanGraphs Def incorporates modern defensive valuation into a broader positional framework.

Therefore, PC1 should not be interpreted as a completely new and independent defensive statistic. A better description would be a consensus advanced-defense index. That is still useful.

In fact, for this particular question, the overlap may be an advantage. If several different defensive systems all point in the same direction, PCA extracts that shared signal and gives it substantial weight.

Chicago is the clearest example. The Cubs do not rank first because one unusual statistic loves their defense. They rank first because virtually every major defensive measure agrees that their defense has been outstanding.

Seattle provides the mirror image. Several independent measurements likewise support the Mariners’ placement near the extreme negative end of PC1.

The most interesting cases may actually be the teams in between. Cincinnati and San Francisco show large disagreements among defensive systems. Pittsburgh has a positive DRS despite poor scores elsewhere. Philadelphia has relatively poor overall PC1 positioning while displaying a much stronger running-game score on PC2.

Those teams deserve additional investigation.

A natural next step also emerges.

Instead of using aggregate measures such as DRS, OAA, and Def, we can construct another PCA using the individual Fielding Run Value components: Throwing, Blocking, Framing, Arm, Range, Infield Double Plays, and First-Base Receiving.

That analysis would answer a different question. This PCA tells us who has been good and who has been bad. A component-level PCA could begin telling us why. For now, though, the broad picture is unusually clear.

Chicago has been the best defensive team in baseball through September 7, and not by a small margin. Arizona forms something of a second tier, followed by a cluster containing St. Louis, Toronto, San Diego, the Dodgers, Atlanta, and Boston.

At the other end, Seattle and the Athletics occupy the weakest part of the defensive landscape.

And perhaps most importantly, the analysis demonstrates why PCA is so useful in baseball analytics.

Defense does not have to be reduced to a debate over which statistic is best. Sometimes the better approach is to let the statistics vote.

 

The Optimization (Short Story Version)

This story has an interesting history.  Many years ago I sat down to write an updated version of what is arguably the greatest short story ever written. If you haven’t read Shirley Jackson’s The Lottery, stop right now and seek it out. It is extraordinary. She wrote it back in 1948, and it still resonates today, perhaps more than ever.

I have extended this story into a novella. I am thinking of expanding it once again into a full novel. I am even considering a trilogy. I have outlined the whole story; I am just waiting until I can find the time to give it the attention it deserves.

Here it the original short story. I kind of like it.

 

The Optimization (Short Story Version)

(Part I)

The morning of June 27th arrived with a brightness that felt almost forced, as if the sun itself had been calibrated to an approved lumen value. Veridian Vale was always clean, but the air on Optimization Day carried a particular sterilized clarity, a faint sting of citrus from the municipal climate diffusers, the mechanical kind of purity that left no room for coincidence. The Steeple, sleek granite, windowless, humming faintly like a throat being cleared, had performed its nightly cleansing cycle hours before dawn. The whole town glistened.

Inside the Larsen household, the walls were already awake, surfacing a slow parade of data: household cohesion score (92), carbon impact (8% below local mean), academic projections (Lily trending upward, Noah stable but variable). Ben Larsen reviewed all of it the way some men once read their newspaper. He stood with his hands clasped behind him, shoulders square, wearing the expression of a man who had long ago learned to keep his inner life trimmed and supervised.

“It’ll be quick,” Ben said as the countdown timer appeared in one corner of the wall-display. “Routine. Nothing to stress about.”

Anya forced a smile that twitched at the edge. She’d been rearranging the same three ceramic tokens on the kitchen counter for nearly ten minutes: a sun, a sprout, a stylized home. Gifts from the Community Uplift Office. Objects intended to inspire unity, gratitude, and reduced cortisol levels.

“I know,” she murmured, though her hands betrayed her, tight and restless, a slight tremor of fear under the skin. “I know, I know.”

Across the room, Lily, twelve, perched on the gel-couch with her legs tucked beneath her, watching Noah instead of the wall. Her brother sat beside her, chin down, fingers fidgeting with the sleeve of his shirt. At sixteen, Noah had the stillness of someone bracing for impact. His hair stuck up in the back (he had slept poorly), and he kept touching the pocket where his school device usually rested. They’d taken it from him last night for “syncing.”

Not unusual. Nothing to worry about.

Except he was worrying. The Larsens could all sense it.

It had started with a moment the previous week, a nothing moment, but it clung to Ben’s memory like a burr. Noah had been at the dining table, hunched over his personal slate, sketching lines with a stylus and a kind of quiet intensity. Ben had walked past, glimpsing only a few sweeping arcs and geometric twists.

“What’s that?” Ben had asked.

“Just something I’m working on.”

“What kind of something?”

Noah hesitated. “A design. For a glider. A real one, not a sim-template. I wanted to mock it up in 3D.”

“All files go to the cloud,” Ben reminded him gently, though his voice carried the crispness of policy, not warmth.

“Yeah,” Noah had said. “I just wasn’t finished.”

Ben hadn’t thought to worry then. Not out loud. But in Veridian Vale, even creativity needed to be tidy and archived.

Now the memory felt like a bruise.

The household display flashed white.

10:00. Optimization Cycle Initiated. Please Connect.

A soft chime bloomed in the air, and all four profile icons lit up: Ben (Senior Data Architect), Anya (Community Health Liaison), Lily (Restricted Juvenile Mode), Noah (Adult Profile, New Status).

“Here we go,” Ben said, steady. Practiced.

But something in the air shifted. A thickness. A waiting.

***

The Night Before

Later, Ben would replay the preceding night as if it held clues he’d missed.

They’d eaten dinner early: nutrient-balanced trays, pale greens, eco-protein, and a small square of community-approved dessert that tasted faintly of almonds. The screens around the neighborhood had pulsed with reminders: Prepare for your Annual Optimization Review. Ensure all devices are synced and charged. Maintain a calm environment. Trust The Guardian.

The phrase hovered in the town like a mantra no one had quite agreed to, but everyone obeyed.

After dinner, Noah had retreated to his room. Ben passed by once and saw light glinting under the door, not the soft blue of standard-issue screens but the warmer glow of his private slate. Ben paused, listening. No voices, no illicit calls, no music sourced from out-of-network channels. Just the scratch of a stylus on digital paper.

“Noah?” Ben knocked lightly. “Everything synced?”

A delay. Small but detectable.

“Yeah, Dad. It’s all good.”

“You handed over your school device?”

“Yep.”

“You uploaded your creative files?”

Silence.

Then: “Most of them.”

“Most isn’t all.” Ben kept his tone even; gentle, but unmistakably directive. “Everything goes to cloud storage before The Review. You know that.”

Noah exhaled, a sigh pulled through his teeth. “I will. I said I will.”

Ben almost stayed. Almost asked What are you hiding? Why does it matter so much? But the rules were simple: trust The System, trust The Process, trust The Guardian.

He walked away with a knot in his stomach, which he tried (unsuccessfully) to ignore.

***

Return to the Present

In the virtual Town Square, beamed into every household’s display and every citizen’s neural implant, 300 households materialized as avatars. Perfect grass. Perfect sky. Perfect order. A place too symmetrical to be anything but artificial.

The Larsen family appeared near the front, their avatars aligned, hands at their sides. Noah’s was a recent scan, shoulders slightly slumped, eyes too serious.

“Hello, Veridian Vale,” Mayor Griffiths said from the pedestal at the center of the green. Her avatar sparkled faintly, airbrushed by municipal protocols. Even her hair had a more mathematically ideal bounce than in person. “Thank you for your presence on this sacred civic day. Let us begin with gratitude.”

A ripple moved through the crowd, habit dressed as devotion.

“Thank The Guardian,” hundreds of voices murmured.

Ben joined in without thinking. Anya did too, though her voice cracked.

Only Noah stayed silent.

The Mayor continued, “The Optimization Cycle ensures Harmony, Safety, and Peak Efficiency for all citizens. Together, we participate in the ongoing purification of our community networks; identifying anomalies, celebrating strengths, and preserving collective well-being.”

Old Man Hendricks’ avatar wavered in the back row. A glitch. He looked smaller this year. Dimmer.

Neighborly whispers, private chat streams, flickered at the edges of consciousness:

He’s failing.
His health data is terrible.
Why didn’t he adjust his diet metrics?

The Review began, household by household.

Chen. Exemplary.
Rourke. Moderate variance in sleep cycles. Acceptable.
Hendricks, M. Declining metrics. Optimization Note: Pending.

A hard swallow echoed faintly across the digital silence. Hendricks spoke, voice filtered, tremulous.

“I updated my medication logs…those scans were private…”

“The Guardian incorporates all consented metrics,” the Mayor replied, her smile tightening at the corners. “Please proceed.”

No more was said. But everyone heard the unspoken truth:
He might not survive this year.

The tension was palpable under the simulation’s perfect sunlight.

It swept next to the Larsen household.

Larsen Household. Collective Efficacy: 92%. Minor anomalies detected.

Anya’s hand found Ben’s. Squeezed just a little too hard.

“Just noise,” Ben said. “Statistical clatter.”

But the data lit up again.

Individual Analysis Required.

Ben’s icon: Stable.
Anya’s: Stable.
Lily’s: Developmentally Stable.

Then Noah.

His profile flickered, not glitching, but hesitating, as if weighing how much truth to reveal.

Larsen, N.
Academic Output: Optimal.
Social Connectivity: Low-Variance.
Creative Activity: Elevated.
Digital Consumption: Anomaly Detected.
Off-network creative files (.psd, .stl) detected on local drive.
Synchronicity Compliance: Below Threshold.
Dissent-Probability Index: 0.06.

A contained murmur rippled through The Square. Nothing more than a subtle expression of disbelief and fear, but also relief that it wasn’t them.

Noah’s avatar froze, suspended in assessment.

Ben felt his stomach drop open.

“That’s wrong,” he said aloud, voice sharper than he intended. “That’s a misclassification. They’re just design files.”

“Dad…” Noah whispered.

“You’re just creative,” Ben insisted, too loudly. “That’s not dissent. That’s…”

But the platform had already changed.

Noah’s school ID photo displayed like a public offering.
Beneath it, in the cold, neutral text of The System:

SELECTED FOR REINTEGRATION.

(Part II)

The word REINTEGRATION hung in the simulated air like a blade suspended by digital thread. It glowed a soft, indifferent blue, the same shade used for household energy reports and school lunch menus. That softness made it worse, a gentle color for a violent verdict.

Anya gasped, the sound small and strangled. Lily clutched her mother’s sleeve.

Noah stared at the display as if staring long enough might change it.

“What does it mean?” he asked, his voice thin, barely audible. “It’s…it’s just the retreat, right? The silent retreat?”

The town had carefully cultivated that lie for years: Reintegration as an off-grid sabbatical, a year of peaceful contemplation, a warm little myth no one dared interrogate. No citizen ever returned to confirm it. Their digital trails went dark, their homes reassigned, their belongings cataloged for redistribution.

But the official line remained, repeated every year like folklore rewritten by committee:
They are reintegrated into The Greater Harmony.
No further details required.

Ben stepped forward in the virtual Square, his avatar half-lit by the swirling data streams around the dais.

“This is a mistake,” he said, projecting authority he no longer felt. “We need a re-evaluation. The anomaly detection is misaligned. He’s not a threat, he’s a child.”

The Mayor’s avatar did something subtle then (barely perceptible), but tangible. She inhaled. A small, constrained breath. Fear, maybe. Or grief. Or irritation at a system she, too, was shackled to. Then she smoothed her expression into the state-sanctioned empathy smile.

“Ben,” she said, her voice warm but hollow, “The Guardian’s process is thorough. You, of all people, understand the precision of the architecture. False positives are statistically negligible.”

Ben’s jaw tightened. “A dissent probability of point-zero-six? That’s rounding error.”

“Even small deviations must be addressed for the health of the whole,” the Mayor said softly. “You know this. Everyone knows this.”

The Square was quiet, 300 households holding their breath and pretending they weren’t. In private chat streams, whispers exploded into viral threads.

Unfortunate.
Such a nice boy.
But anomalies spread.
And the family’s score would’ve plummeted if they didn’t… comply.

Ben heard none of it, but he could feel it. The pressure was physical, a tightening of invisible fingers around the throat of his household.

“Please,” Anya said, her voice cracking. “Please, there must be a manual override, a petition process…something…”

“There is no override,” the Mayor said. “Once The Guardian identifies a Reintegration candidate, the decision is final.”

Noah’s avatar didn’t move, but in the real world, he was shaking, small tremors in his arms, his shoulders, the kind of shaking that begins at the core.

“Mom?” he whispered. “I’m not…I didn’t…”

“I know,” she said, pulling him close even though they were only holograms here. “I know, baby.”

But knowing didn’t matter.

The Guardian had spoken.

***

Before the Stoning

In the Larsen living room, reality snapped back as The Metaspace dissolved. The wall-display remained active, pulsing faintly with Noah’s profile summary and a countdown clock:

Awaiting Community Consensus (14 seconds)

A soft ping, gentle and courteous, echoed through the room.

Ben looked down at his own wrist implant.

Guardian Request:
Confirm Network Consensus for Reintegration of User: N_Larsen.
To maintain your household’s trust status, please acknowledge.

Below the text: CONFIRM or REQUEST CLARIFICATION.

He knew what clarification meant. Anomaly stacking. Suspicion by association. A mark placed on the entire household, one that might not be reversible.

Across Veridian Vale, every device lit up in perfect synchrony. Residents saw the same prompt. Their hearts beat in the same fearful rhythm. Their fingers hovered over the same options.

To save themselves, they had to condemn him.

The digital stoning had begun.

Anya’s face was pale, almost gray. “Ben,” she whispered, “what do we do?”

Ben’s throat felt scraped raw. He looked at his son, sixteen, trembling, eyes wide with an animal kind of fear.

“Dad,” Noah whispered. “Please don’t. I didn’t do anything.”

And it was true. He hadn’t.

But truth didn’t move The Guardian.

Ben’s thumb hovered over REQUEST CLARIFICATION.
His pulse spiked. The implant vibrated a warning: elevated stress detected.

Lily began to cry softly, muffled, as if sound itself might attract The Guardian’s attention.

Ben closed his eyes.

Across town, fingers pressed CONFIRM.
Soft tones chimed in living rooms like polite applause.

One by one, neighbors sealed Noah’s fate.

Ben’s thumb trembled. He lowered his hand.
Anya sagged against the counter, relief and shame sliding together across her face.

The household remained in good standing.

Noah was not.

***

The Erasure

A secondary display on the wall lit up automatically, beginning the public dissolution of Noah Larsen.

It was both meticulous and indifferent, like an accountant closing an account.

School ID: SUSPENDED.
Transit Pass: REVOKED.
Bank Account: FROZEN.
Social Profiles: DEACTIVATED.
Health Records: ARCHIVED (LOCKED).
Household Biometric Access: REMOVED.
Device Authentication: CANCELLED.

Line by line, Noah vanished.

He choked out a sound (half-sob, half-breath), but he didn’t fight. Not outwardly. The Guardian absorbed resistance the way a black hole absorbed light: nothing escaped.

Ben reached for him, but the wrist implant buzzed—physical interference with Reintegration Protocol will result in status review. Even touch was regulated today.

Noah stepped backward instead, pressing himself into the corner of the room as if he might merge with the wall and disappear before The System erased him.

***

The Vehicle Arrives

Through the glass front wall, a white autonomous vehicle slid silently to the curb; a minimalist pod with no windows except a smooth, opaque front panel. It opened without a sound.

Two Reintegration Specialists emerged. They were humanoid, but not human. Soft synthetic skin, gentle features, and voices tuned to the optimum frequency for reducing panic.

“Noah Larsen,” the first said, its tone soothing as warm water. “Please come with us. Your Reintegration journey awaits.”

Noah didn’t move.

Ben stepped in front of him on instinct, but Anya grabbed his arm, a sudden, crushing grip.

“Don’t,” she whispered, eyes wild with terror. “Please, Ben, don’t. We can’t. We can’t lose Lily too.”

Ben froze.

He imagined pressing REQUEST CLARIFICATION.
He imagined pulling Noah back, slamming the door, barricading the house.

But The Guardian watched everything.

Noah looked between his parents, confusion giving way to betrayal.

“Dad?” he asked again, voice cracking. “Mom?”

Anya crumpled, burying her face in her hands.

Ben’s heart thrashed. He tried to speak, tried to find words that might soften this moment, something like I love you, or I’m sorry, or even I failed you.

But nothing came.

The Specialists stepped forward, their movements gentle, inevitable.

Noah didn’t struggle. He simply let his body move where guided, like someone who had run out of choices.

As they led him toward the vehicle, he turned back. His face wasn’t angry. It wasn’t even afraid.

It was unreadable. As if he had already been erased.

The pod closed.
It rolled away in silence.
No destination displayed.

***

Completion

A final chime sounded through the Larsen home:

Optimization Cycle Complete.
Veridian Vale Harmony Score Updated: 98.7%.
Thank you for your participation.

The climate system released a faint lavender scent into the air, the fragrance of civic compliance and emotional regulation.

The family portrait on the wall flickered.

Noah’s place disappeared. The background stitched itself seamlessly, as if he had never existed at all.

Anya collapsed onto the gel-couch. Lily buried her face in her mother’s side.

Ben stared at the blank space where his son’s avatar had been.

The house was silent, but not empty.

Silence had weight. Silence had shape.

Outside, Old Man Hendricks sat alone in his small, dim living room. His own data-spike forgotten for now, he muttered to no one:

“Before the Guardian, we had crime, inefficiency, waste. We weren’t optimized.”

He repeated it, softer this time, as if trying to reassure himself.

“We weren’t optimized.”

And the town, relieved, returned to its quiet, perfect, data-driven day.

(Part III)

The lavender scent lingered long after the notification dissolved, settling into the corners of the Larsen home like a fog engineered to suppress grief. The municipal guidelines described it as Mood Harmonizer #4, a patented blend calibrated to “mitigate emotional disequilibrium following civic participation.”

But no amount of scent could make the house feel whole.

Ben didn’t move for a long time. He stood in front of the display, staring through the empty space where Noah’s avatar had been, as if the right angle, the right squint, might reveal a ghost outline. The System had been thorough; it always was. Not even residual pixels remained.

Anya’s hand slid into his, cold and trembling. She didn’t speak. There was no script for this part, no comforting instructions, no municipal packet on “Coping After Reintegration of a Household Member.” That sort of literature might imply that the process was traumatizing, and the Guardian did not permit negative framing.

Lily had fallen asleep on the couch, cheeks still damp. Children adjusted fastest; that was what the training modules claimed. Their minds were more elastic. They absorbed loss like water into sand.

Ben wished he could believe it.

He sat beside her and brushed a strand of hair from her forehead. She didn’t stir.

“I should’ve pressed Clarification,” he said finally, his voice so low it barely qualified as sound.

Anya didn’t look at him. Her eyes were locked on the front door, still closed, still silent, but feeling like an open wound.

“You saw the warning,” she whispered. “They would’ve flagged us. Ben… they would’ve taken her too.”

He knew. Of course, he knew. The Guardian rarely acted on single anomalies. It was pattern recognition that mattered. Association. Contagion. A household that challenged a Reintegration decision risked becoming a cluster needing adjustment.

“We could’ve tried,” Ben said, but the words rang hollow.

“And then what?” Anya’s voice broke. “What would’ve happened? Both kids gone? Or all of us? What then?”

Ben looked down at his hands. They didn’t feel like his. “He’s not dangerous. He’s not even rebellious. He’s…”

“Was,” Anya whispered, and then she closed her eyes as if the word itself hurt.

***

The Void Noah Left

In the corner of the living room, Noah’s personal slate sat on the floor, propped against the wall where he’d dropped it earlier. The System had already tried to auto-wipe it; a thin progress bar flickered beneath an error message:

Local Encryption Detected.
Manual Override Required.
Reintegration Protocol supersedes data privacy.

Ben swallowed hard.

“Did you know he encrypted his files?” he asked.

Anya shook her head. “He said they were just designs.”

Ben picked up the slate. It vibrated faintly, rejected by his biometric signature.

“Why would he encrypt them?” Anya asked, barely above a whisper.

“I don’t know,” Ben said. “Maybe he was proud of them. Maybe he didn’t want others to copy them. Maybe he wanted something to be his.”

As soon as he said it, he understood.

In Veridian Vale, nothing belonged to individuals. Creativity was communal. Innovation was monitored, quantified, and redistributed. There was no privacy, only shared metrics. Only transparency.

Maybe a sixteen-year-old had simply wanted a corner of the world that The Guardian didn’t supervise.

Maybe that was enough to be declared dangerous.

***

Across Town

While the Larsens sat in the quiet aftermath, the town hummed with cautious relief.

At the Chen household, a celebratory drink capsule hissed open, ginger turmeric, optimized for cardiovascular longevity. “Such a shame about the Larsen boy,” Mrs. Chen said, not sounding particularly ashamed. Her husband nodded with a frown that didn’t reach his eyes. “But The Guardian sees farther than we do.”

At the Rourke residence, the family gathered around the dinner table, voices low. “They should’ve synced those files,” Mr. Rourke muttered. “Everyone knows better.” His daughter, Mae, stared at her plate. She had once worked on a project with Noah, an art assignment. She remembered he had been kind. Quiet. Smart. She remembered thinking he made the virtual world feel a little less artificial.

She said nothing.

In Old Man Hendricks’ home, he sat in his reclining pod, hands trembling against the armrests. He’d expected his name on the stage this year. Maybe he was relieved. Perhaps he was horrified that he was relieved.

“Waste,” he muttered to the empty room. “Before The Guardian, there was waste.”

He said it again and again, a mantra meant to convince himself that efficiency was worth the cost.

But he didn’t sound convinced.

***

The Night After

The Larsens didn’t sleep.

Ben lay awake listening to the house breathe, the soft hiss of the climate vents, the faint electrical hum behind the walls, the pulsing glow of the utility panel. All of it woven into a single low-frequency reminder that the Guardian was always awake.

At 2:14 a.m., his wrist implant buzzed.

Household Stress Index Elevated.
Consider listening to a Harmonizer track.
Recommended: “Waves of Unity, Track 6.”

He silenced the suggestion.

When he rose to use the bathroom, he saw Noah’s door open, dark, empty. The System had already begun a baseline refresh. His sheets had been sanitized. His mattress recalibrated. His posters were removed from the wall and recycled into digital nothingness. The room smelled of antiseptic lavender.

Ben stepped inside.

Everything familiar had been stripped, wiped, streamlined.

He touched the wall panel beside the bed. It flickered, then displayed:

Occupant ID: NONE
Room Available for Reassignment
Estimated Reallocation: 5 days

The speed of it made him nauseous.

He sat on the edge of the bed, the gel-cushion adjusting to his weight with an eager efficiency that felt obscene.

He saw Noah as he’d last looked at him, eyes wide, not angry, not pleading, just…absent, already drifting from himself even before The Specialists reached him.

Ben pressed his palms to his eyes until he saw stars.

***

The Following Morning

The town woke to another perfect day, temperature regulated to 22°C, humidity balanced, sunlight filtered through particulate screens.

Ben stood at the kitchen counter, nursing a cup of nutrient coffee. The wall display showed the usual morning metrics, but now there was a hollow rectangle where Noah’s academic updates had scrolled.

Anya set a plate in front of Lily, who stared into her cereal, unmoving.

“You need to eat,” Anya said.

“I’m not hungry.”

“You need to eat,” Anya repeated, voice too firm, too brittle.

Lily lifted her spoon with a trembling hand.

Then, a chime at the door. Ben froze; no one visited unannounced.

He exchanged a look with Anya. She stepped behind Lily protectively.

The door slid open.

A Reintegration Specialist stood on the threshold, one of the same models that had taken Noah, features molded into permanent calm.

“Good morning,” it said. “This is a wellness follow-up. The Guardian has detected abnormal stress indicators in this household.”

Ben’s pulse spiked. “We’re fine.”

The android’s gaze did not change. “To maintain optimal cohesion, The Guardian recommends a Resilience Consultation.”

“We’re fine,” Ben repeated.

The Specialist stepped forward, just enough to test boundaries. “Refusal will be logged.”

Anya cleared her throat. “We accept.”

Ben turned toward her sharply. “Anya…”

“We accept,” she said again, louder this time. Her eyes were pleading, not at The Specialist, but at Ben. Do not fight this. Do not risk us further.

The Specialist nodded. “Your consultation is scheduled for 14:30. A reminder will be sent.”

Then it departed, no threat, no weapon, no raised voice. All parties realized the inevitability.

Ben leaned against the counter after the door slid shut. His hands shook.

“They’re watching us,” he said.

“They always were,” Anya whispered.

***

A Glitch in The System

After The Specialist left, Ben returned to Noah’s encrypted slate. He tried again to access it. Again, it rejected him.

But this time he noticed something new, an icon flickering in the corner of the display. Barely there, a thin crack of mismatched color.

A glitch.

Noah had written code before, school assignments, minor sandbox games, small creative widgets. Nothing dangerous. Nothing The Guardian would normally care about.

But maybe this time he had built something more.

Something The System couldn’t immediately erase.

Ben tapped the flickering corner. The slate blinked, then opened a single file—its encryption bypassed only long enough to display a single line of text, rendered in Noah’s handwriting:

If someone sees this, it means I wasn’t careful enough.
But it also means I was right. There’s something wrong in The System.
I think The Guardian is hiding—

The line cut off abruptly.

The screen went black.

A new message replaced it:

Unauthorized Access Attempt Detected.
Device Memory Purged.
Thank you for maintaining community safety.

Ben stared at the blank slate.

“What was he trying to tell us?” he whispered.

No one answered.

***

Closing

In Veridian Vale, the day unfolded without incident. The climate system misted jasmine into the air. Children walked to school in neatly monitored lines.
Drones drifted overhead like guardian insects. Neighbors exchanged polite greetings calibrated to optimal decibel ranges.

Harmony. Order. Efficiency.

Noah Larsen was gone, absorbed into the immaculate machinery of civic perfection.

And in the Larsen home, beneath the lavender and jasmine, beneath the silence and the screens, something new began to take root.

A question. A dangerous one.

Not spoken aloud. Not recorded. (Not yet).

But living, growing, waiting.

The kind of question that could fracture a perfect system.

The kind of question The Guardian feared most.

 

 

The Double Life of Véronique

The Double Life of Véronique

Krzysztof Kieslowski’s The Double Life of Véronique is not merely a great film; it is an extraordinary sensory experience. It is a work of profound beauty and melancholy, a meditation on fate, identity, and the invisible threads that connect us all, and it remains, at least to me, a masterpiece in the landscape of cinema.

The plot is deceptively simple and exceedingly ambiguous. Two women, both played with luminous sensitivity by the great Irène Jacob, live separate lives on opposite sides of Europe: Weronika in Krakow, Poland, and Véronique in Paris, France. They are identical in appearance, share a gift for music, and are haunted by a vague, unexplained (and mystical) feeling that they are not alone. Their lives are entangled in a metaphysical bond that neither fully understands. In a fleeting moment, their paths almost cross in a Krakow square, but the connection is missed. When one makes the ultimate sacrifice for her art, the other feels a sudden, inexplicable loss that will change the course of her life.

The genius of the film lies not in its narrative, but in how it is told. Kieslowski (a full-fledged genius), working with his regular cinematographer, Slawomir Idziak, creates a world of light, reflections, and distortions that visually represents the film’s themes. Windows, mirrors, camera lenses, and even glass spheres are used to create a sense of constant doubling, fragmentation, and entanglement. The golden filters saturate the screen, creating a world of warmth that is at once dreamlike and fragile, perfectly mirroring the protagonists’ emotional states.

Central to the film is Irène Jacob, who rightfully won the Best Actress award at Cannes. She is in almost every scene, bringing two distinct characters to life while subtly suggesting their shared soul. She is not just an actor; she is the film’s beating heart, its symbol of grace and vulnerability.

Zbigniew Preisner’s haunting score is an equally crucial element, an achingly beautiful presence that underscores the film’s spiritual and emotional weight. The music is a character in itself, its melodies weaving a spell of nostalgia and loss that lingers long after the credits roll. His score is as elusive as it is beautiful.

So, what is the deal here? Why am I writing about this movie? There is a reason, which I believe is a pretty good one. I put off watching this film for quite some time because I wanted to have something to look forward to. A few days ago, as I sat in my chair suffering through a severe arthritis flare in my ankle, I decided to give in and watch. This is what happened.

Movies are stories, and the creators hope viewers are drawn in and absorbed by the screen. I must admit that the screen had my full attention, but I kept getting lost in the story. Guesses? Any ideas how and why?

I find Irène Jacob so beautiful that she continually distracted me. I had the same problem with Three Colours: Red the first few times I watched it. Apparently, an actress can be so attractive that I cannot follow the story. And that, for whatever it is worth, is my story.

 

 

Which MLB Teams Are Actually Pitching the Best in 2026? (A Team Pitching Study Through September 4, 2026)

ERA is useful. It is also incomplete.

A team can post an excellent ERA because its pitchers dominate hitters, limit hard contact, avoid walks, and miss bats. But a good ERA can also reflect defense, sequencing, favorable outcomes with runners on base, or simply a stretch in which balls have found gloves instead of grass.

The reverse can happen too. A pitching staff can do many of the things we associate with good pitching and still carry an ERA that makes it look merely average.

That distinction becomes particularly interesting when we look across all 30 major-league teams in 2026.

Using FanGraphs team pitching data through September 4, I wanted to answer a slightly different question from the usual one. Instead of asking which teams have allowed the fewest earned runs, I asked:

Which teams appear to be pitching the best underneath those results?

That question leads us to Philadelphia.

But it also leads us to Milwaukee, Cleveland, Atlanta, Arizona, the Yankees, and several teams whose records do not line up neatly with the quality of their pitching.

Measuring the Pitching Process

No single pitching statistic completely describes a staff.

ERA tells us what happened. FIP focuses heavily on strikeouts, walks, hit batters, and home runs. xFIP normalizes the home-run component. SIERA attempts to estimate run prevention while accounting for strikeouts, walks, and batted-ball tendencies. xERA uses Statcast information about the quality of contact allowed.

Then there is K-BB%, one of the cleanest measures of pitcher control.

A staff that strikes out many hitters while walking few is usually doing something right.

I therefore built an Underlying Pitching Process Index using seven measures:

\mathcal{M} = \left\{ \mathrm{xERA}, \mathrm{FIP}^{-}, \mathrm{xFIP}^{-}, \mathrm{SIERA}, \mathrm{K\!-\!BB\%}, \mathrm{Barrel\%}, \mathrm{HardHit\%} \right\}

The first step was to standardize every statistic across the 30 teams.

For team and pitching metric :

z_{i,m} = \frac{ x_{i,m} - \overline{x}_m }{ s_m }

There is an important complication.

Higher K-BB% is good. Higher xERA is not. The same problem applies to FIP-, xFIP-, SIERA, Barrel%, and HardHit%.

I therefore adjusted the direction of each standardized statistic so that higher always means better pitching:

q_{i,m} = d_m z_{i,m}

where:

d_m = \begin{cases} +1, & \mathrm{if\ higher\ values\ are\ better} \\ -1, & \mathrm{if\ lower\ values\ are\ better} \end{cases} q_{i,m} = d_m z_{i,m}

where:

d_m = \begin{cases} +1, & \text{if higher values are better} \\ -1, & \text{if lower values are better} \end{cases}

The seven adjusted z-scores were then averaged:

P_i = \frac{1}{7} \sum_{m \in \mathcal{M}} q_{i,m}

Finally, I converted that score into an index centered on 100, with a standard deviation of 15:

I_i = 100 + 15 \left( \frac{ P_i - \overline{P} }{ s_P } \right)

A score of 100 represents approximately league-average underlying pitching. A score of 115 is about one standard deviation above average.

This is important: the Process Index is not a FanGraphs statistic. It is a composite measure I constructed for this study from FanGraphs data.

And one team separates itself immediately.

Philadelphia Comes Out on Top

Figure 1 ranks all 30 teams using the Underlying Pitching Process Index.

Figure 1. 2026 MLB Team Pitching Through September 4: Underlying Pitching Process Ranking.

The index combines xERA, FIP-, xFIP-, SIERA, K-BB%, Barrel% allowed, and HardHit% allowed. The MLB average is centered at 100.

Philadelphia ranks first with a Process Index of 131.2.

Milwaukee is close behind at 128.6. The Yankees rank third, followed by Boston and the Dodgers.

The top ten are:

Rank Team Process Index
1 Philadelphia 131.2
2 Milwaukee 128.6
3 New York Yankees 121.6
4 Boston 118.6
5 Los Angeles Dodgers 117.4
6 Toronto 110.2
7 Cleveland 109.3
8 New York Mets 108.0
9 Pittsburgh 107.3
10 Detroit 107.3

Philadelphia’s position may initially seem strange.

The Phillies have a 4.00 ERA. Nobody looking only at ERA would identify that as the performance of baseball’s best pitching staff.

The underlying numbers tell a different story.

Philadelphia has a 3.56 xERA, 3.76 FIP, 3.54 xFIP, and 3.50 SIERA. Most impressively, the Phillies have a 17.9% K-BB%, the best figure in baseball through September 4.

That combination is difficult to dismiss.

They miss bats. They limit walks. Their contact profile is excellent. Several statistics designed specifically to isolate pitcher performance show a staff substantially better than its 4.00 ERA.

Milwaukee presents a different case.

The Brewers rank second in underlying process, with an index of 128.6, but their actual ERA is already excellent at 3.50. Their xERA is 3.47.

For Milwaukee, there is very little disagreement between process and outcome.

That distinction becomes important.

ERA and xERA Do Not Always Tell the Same Story

One simple way to measure the difference between observed and expected run prevention is:

\Delta \mathrm{ERA}_i = \mathrm{ERA}_i - \mathrm{xERA}_i

A positive value means the team’s ERA has been worse than its xERA. A negative value means the team has allowed fewer earned runs than its expected ERA would suggest.

Figure 2 shows the relationship directly.

Figure 2. 2026 Team ERA vs. xERA Through September 4.

Teams below the diagonal have an ERA lower than their xERA. Teams above the diagonal have an ERA higher than expected from their Statcast contact profile.

Philadelphia stands out.

For the Phillies:

\Delta \mathrm{ERA}_{\mathrm{PHI}} = 4.00 - 3.56 = 0.44

Their ERA is approximately 0.44 runs higher than their xERA.

That is a meaningful gap over nearly an entire season.

The Athletics are the extreme example of the opposite kind of outcome. Their ERA is 5.47, while their xERA is 4.48:

\Delta \mathrm{ERA}_{\mathrm{ATH}} = 5.47 - 4.48 = 0.99

Oakland’s underlying pitching is not good. The Athletics rank only 24th in the Process Index.

But a 5.47 ERA makes the staff appear even worse than its underlying performance suggests.

Now look below the diagonal.

Arizona has a 4.17 ERA against a 4.72 xERA:

\Delta \mathrm{ERA}_{\mathrm{ARI}} = 4.17 - 4.72 = -0.55

Atlanta is similar:

\Delta \mathrm{ERA}_{\mathrm{ATL}} = 3.57 - 4.04 = -0.47

The Cubs sit at 4.16 versus a 4.62 xERA, while the Yankees have produced an exceptional 3.23 ERA despite a 3.67 xERA.

Those differences do not mean that one statistic is right and the other is wrong.

They tell us that something interesting is happening between underlying pitcher performance and actual runs allowed.

Process Is Not the Same as Results

To examine that distinction more directly, I created a second index.

The Results Index uses ERA-, RA9-WAR, and WPA. Unlike the Process Index, these measures deliberately capture more of what actually happened on the field.

The underlying results score is:

R_i = \frac{1}{3} \left( -z_{i,\mathrm{ERA}^{-}} + z_{i,\mathrm{RA9\!-\!WAR}} + z_{i,\mathrm{WPA}} \right)

It is then converted to the same 100-centered scale:

J_i = 100 + 15 \left( \frac{ R_i - \overline{R} }{ s_R } \right)

Figure 3 compares the two indices.

Figure 3. 2026 Team Pitching: Underlying Process vs. Actual Results.

Teams above the diagonal have obtained better results than their underlying pitching process would predict. Teams below the diagonal have performed worse.

The diagonal is the key.

Teams close to it have received results consistent with the quality of their underlying pitching. Teams far above or below it deserve more attention.

Philadelphia ranks first in process but only eighth in results.

Its Results Index is almost 20 points below its Process Index.

Arizona goes the other way. The Diamondbacks rank only 25th in underlying process, but 11th in results.

Atlanta is another striking example. The Braves rank 12th in process but second in results.

The Cubs rank 26th in process and 17th in results.

These teams have converted their pitching performances into actual run prevention more effectively than the underlying indicators alone would predict.

Philadelphia has not.

Neither have the Athletics, the Seattle Mariners, the San Francisco Giants, or the Mets.

There are many possible reasons. Defense is a factor. So does sequencing. Relievers inherit runners. Pitchers behave differently with runners on base. Ballparks are important as is random variation.

The point is not that all deviation must eventually disappear.

It is that ERA alone completely hides the deviation.

Strikeouts, Walks, and SIERA

One of the strongest relationships in the data appears when we compare K-BB% with SIERA.

Figure 4. 2026 Team Command and Dominance Through September 4.

Teams toward the right strike out more hitters relative to walks. Teams toward the top have lower SIERA.

K-BB% is appealing because it strips pitching down to two outcomes over which the pitcher exercises considerable control.

The calculation is simple:

\mathrm{K\!-\!BB\%} = \mathrm{K\%} - \mathrm{BB\%}

Philadelphia again occupies elite territory.

The Phillies lead MLB at 17.9%. Milwaukee follows at 17.3%. The Dodgers are at 16.5%, with the Yankees and Cleveland both around 16%.

This is one reason Cleveland deserves more attention.

The Guardians are not surviving through a suspiciously low ERA or an unusually favorable gap between ERA and expected statistics. They rank fifth in SIERA and fifth in K-BB%.

That is a much stronger foundation.

Cleveland’s Pitching Is Not the Problem

Cleveland entered September 5 with a 72-70 record.

The Guardians’ pitching numbers, however, look like those of a considerably better team.

Cleveland ranks seventh in ERA, eighth in xERA, sixth in FIP, sixth in xFIP, fifth in SIERA, fifth in K-BB%, seventh in pitching WAR, seventh in the Underlying Pitching Process Index, and seventh in the Results Index.

That consistency is important.

Cleveland is not seventh because one unusual statistic dragged the composite upward. Almost every important run-prevention and defense-independent measure puts the Guardians somewhere in the same general neighborhood.

Their ERA is 3.75.

Their xERA is 3.90.

Their FIP is 3.86, xFIP is 3.88, and SIERA is 3.75.

There is, however, one weakness.

Cleveland ranks only around 21st in Barrel% allowed and 22nd in HardHit% allowed. Opponents have made better contact against the Guardians than the overall pitching rankings might imply.

They have compensated for that weakness largely through strikeouts and control.

That creates a somewhat unusual pitching profile. Cleveland is not especially dominant at suppressing every kind of hard contact. Still, the Guardians prevent enough balls from being put into play in the first place and avoid giving away enough free bases to remain an excellent overall staff.

If the question is why Cleveland has been only a little above .500, the team’s pitching data strongly suggests looking elsewhere.

How Much Does Pitching Explain Winning?

Of course, a baseball team does not win with pitching alone.

To measure the association between underlying pitching quality and the standings, I compared the Process Index with team winning percentage.

Winning percentage is:

\mathrm{WinPct}_i = \frac{ W_i }{ W_i + L_i }

I then estimated a simple linear regression:

\mathrm{WinPct}_i = \alpha + \beta I_i + \varepsilon_i

The resulting relationship appears in Figure 5.

Figure 5. 2026 Underlying Pitching Quality vs. Team Winning Percentage.

The regression compares each team’s Process Index with its winning percentage through September 4.

The coefficient of determination is:

R^2 = 0.314

Approximately 31.4% of the cross-team variation in winning percentage is associated with variation in the Underlying Pitching Process Index in this 2026 snapshot.

Pitching clearly is essential, but nearly 69% of the variation remains elsewhere.

The Cubs are 80-62 despite ranking 26th in underlying pitching process. Atlanta is 84-57 while ranking only 12th. Cleveland is 72-70 despite ranking seventh.

A team can compensate for mediocre pitching.

It can also waste very good pitching.

Philadelphia May Be Better Than Its ERA

The Phillies may be the most important example in the entire study.

Their 4.00 ERA does not look dominant. If we stopped there, we would never consider Philadelphia the best pitching team in baseball.

Yet the deeper indicators repeatedly return to the same conclusion.

Philadelphia ranks first in the Process Index. The Phillies lead baseball in K-BB%. Their xERA is 3.56. Their xFIP is 3.54. Their SIERA is 3.50.

Their actual ERA is the outlier.

That does not guarantee that the ERA will suddenly collapse toward 3.50. Baseball statistics do not work like a mechanical spring that must snap back to equilibrium.

But if I were trying to determine which pitching staff I trusted going forward, I would place more weight on the cluster of underlying measures than on ERA alone.

Philadelphia’s process has been elite.

Milwaukee Has the Cleaner Case

Milwaukee requires less explanation.

The Brewers rank second in the Process Index at 128.6 and have a 3.50 ERA against a 3.47 xERA.

There is almost no gap.

Their process says they are excellent. Their results say they are excellent. Their 88-54 record agrees.

Sometimes the complicated analysis confirms the obvious.

That is useful too.

Milwaukee does not need a regression argument or a discussion of sequencing to explain its success. The Brewers have simply pitched extremely well.

The Yankees Have Turned Good Pitching Into Great Results

The Yankees provide another variation.

New York ranks third in underlying process but first in the Results Index.

Their 3.23 ERA is substantially better than their 3.67 xERA:

\Delta \mathrm{ERA}_{\mathrm{NYY}} = 3.23 - 3.67 = -0.44

The underlying staff is genuinely good. This is not a weak pitching team disguising itself behind a low ERA.

But the results have been even better than the already strong process.

That distinction is noteworthy.

The Yankees do not appear to be a mirage. They appear to be an excellent pitching staff that has also benefited from a gap between underlying performance and runs actually allowed.

Arizona Is More Difficult to Trust

Arizona presents a much different problem.

The Diamondbacks’ 4.17 ERA is not particularly impressive on its own, but it looks much better compared to their 4.72 xERA.

Their Process Index ranks only 25th in baseball.

Their Results Index ranks 11th.

That is one of the largest process-to-results gaps in the league.

Perhaps Arizona can continue converting that underlying performance into acceptable run prevention. There may be legitimate reasons for part of the difference.

Still, if the goal is prediction rather than description, I would be considerably more cautious about Arizona than Philadelphia.

The Phillies have poor results relative to a strong process.

Arizona has strong results relative to a poor process.

Those are not equivalent situations.

Oakland Shows How Ugly Results Can Become

At the bottom end, Oakland offers an equally useful lesson.

The Athletics rank 24th in underlying pitching process. That is bad, but not the worst in baseball.

Their Results Index ranks 30th.

Their 5.47 ERA is almost one full run higher than their 4.48 xERA.

The underlying numbers therefore suggest two conclusions at once.

Oakland has not pitched well.

Oakland has also gotten even worse results than its mediocre pitching would normally lead us to expect.

Both statements can be true.

What the Rankings Really Tell Us

This study is not intended to replace ERA with another single magic number.

In fact, that would defeat the purpose.

The more interesting lesson is that pitching has several layers.

There is the process: missing bats, limiting walks, controlling contact, suppressing barrels, and producing outcomes that FIP, xFIP, SIERA, and xERA consider sustainable.

Then there are the results.

Usually they move together.

Sometimes they do not.

Philadelphia has arguably produced the strongest underlying pitching process in baseball, yet received only the eighth-best composite results.

Milwaukee has been excellent by both standards.

The Yankees have turned excellent underlying pitching into even better results. Atlanta and Arizona have substantially outperformed what their process metrics would lead us to expect.

And Cleveland?

Cleveland may be the most revealing team of all.

The Guardians are barely above .500, yet almost every pitching measure says the staff belongs among the top quarter of baseball. Their Process Index ranks seventh. Their Results Index also ranks seventh.

There is little evidence that Cleveland’s pitching has betrayed the team.

If anything, it has kept the Guardians afloat.

Final Thoughts

ERA remains one of baseball’s most intuitive statistics because it measures something real. Runs crossed the plate, and those runs counted.

But describing what happened is not always the same as understanding why it happened.

That is where the deeper statistics become valuable.

Through September 4, Philadelphia appears to have baseball’s strongest underlying team pitching process. Milwaukee is close behind and has translated that process into superior run prevention. The Yankees have done even better in converting strong pitching into actual results.

Cleveland quietly belongs in the next group.

Meanwhile, Atlanta, Arizona, and the Cubs remind us that actual run prevention can exceed expectations based on underlying statistics, sometimes by a considerable margin.

None of those observations requires us to choose between ERA and advanced metrics.

We can use both.

ERA tells us what the scoreboard recorded.

The underlying statistics help us understand how the pitchers got there and perhaps where they are going next.

That difference is where the interesting part begins.

 

Who Is Producing the Most Offense in Baseball in 2026 Through Early September?

We tend to talk about hitting leaderboards as though they answer a simple question. Who has been the best hitter?

It is not quite that simple. A hitter can produce enormous value through power, another through reaching base, another through a combination of contact and baserunning. Rate statistics can identify extraordinary performance while overlooking playing time. Counting statistics reward accumulation but can obscure efficiency. WAR goes even further, mixing offense with defense and positional value.

For this study, I wanted to ask a narrower question: Who has actually produced the most offensive value in Major League Baseball so far in 2026?

I used my September 5 FanGraphs downloads, containing 142 hitters and a wide range of traditional, advanced, Statcast, baserunning, and value statistics. The results reveal a clear No. 1. They also reveal something more interesting underneath.

Measuring offensive production

The primary statistic for the study is FanGraphs Offense. In the data, it has a particularly useful decomposition:

\mathrm{Offense}_i = \mathrm{Batting}_i + \mathrm{BaseRunning}_i

That makes it attractive for this question.

Defense is absent. Positional adjustment is absent. We are measuring what a player has contributed when his team is at the plate or when he is running the bases.

WAR is still useful, but it answers a different question. Here I wanted to isolate offense.

The initial leaderboard is striking.

Yordan Alvarez leads baseball with 52.31 offensive runs above average. Pete Crow-Armstrong is very close behind at 49.45.

Then comes a cliff.

James Wood is third at 33.68, followed by Shohei Ohtani at 32.85 and Randy Arozarena at 30.61. The difference between Alvarez and Crow-Armstrong is only 2.86 runs. The difference between Crow-Armstrong and Wood is nearly 15.8 runs.

So, at least by this measure, 2026 has developed a clear top tier of two.

Alvarez and Crow-Armstrong arrived there differently

The decomposition is useful because Alvarez and Crow-Armstrong have not constructed their offensive value in the same way.

For Alvarez:

\mathrm{Offense}_{\mathrm{Alvarez}} = 57.01 - 4.70 = 52.31

His bat has been so productive that he can lose almost five runs through baserunning and still lead everyone.

Crow-Armstrong presents almost the opposite profile:

\mathrm{Offense}_{\mathrm{CrowArmstrong}} = 43.77 + 5.68 = 49.45

His batting contribution is excellent, but his legs add another 5.68 runs.

That distinction can be seen clearly when batting and baserunning are separated.

Alvarez is sitting far to the right because of his extraordinary batting value, but below zero in baserunning. Crow-Armstrong combines one of baseball’s best bats with strongly positive running value.

James Wood and Ohtani fall somewhere between those extremes.

There is no single recipe.

Rate production tells almost the same story

A natural objection to using total offensive runs is that playing time matters. A player can accumulate more value simply because he has received more plate appearances.

That is where wRC+ becomes useful.

Across the 142 hitters in the dataset, FanGraphs Offense and wRC+ have a correlation of approximately:

r = 0.970

A simple regression gives:

\widehat{\mathrm{Offense}}_i = -60.98 + 0.618 \left( \mathrm{wRC}^{+}_i \right)

with:

R^2 = 0.941

That is an exceptionally tight relationship.

It should not be interpreted as independent validation, since Offense and wRC+ are built from related offensive information. What it does show is that differences in playing time and baserunning have not radically reordered the best hitters. The strongest rate producers are generally producing the most total offensive value as well.

And once again, Alvarez sits at the top.

His 180.5 wRC+ is the best in the dataset. Crow-Armstrong follows at 157.6, while Willson Contreras, James Wood, Bryce Harper, Randy Arozarena, Junior Caminero, and Ohtani form the next group.

An approximately 181 wRC+ means that Alvarez has created runs at roughly 81 percent above league average, after the adjustments built into wRC+.

That is a remarkable offensive season.

But what should have happened?

Observed production is only half of the story.

Statcast gives us another way to look at these hitters. Instead of merely asking what happened, xwOBA asks what we would expect from the quality of the contact, along with the other inputs incorporated into the statistic.

I calculated the difference as:

\Delta \mathrm{wOBA}_i = \mathrm{wOBA}_i - \mathrm{xwOBA}_i

A positive value means the player’s actual wOBA has exceeded his xwOBA. A negative value indicates that his expected production is higher than his observed production.

The comparison changes the story.

Alvarez has a .4289 wOBA, which is already the best in the sample.

His xwOBA is .4496.

\Delta \mathrm{wOBA}_{\mathrm{Alvarez}} = 0.4289 - 0.4496 = -0.0207

In other words, the underlying contact data do not suggest that Alvarez has been getting lucky.

They suggest the opposite.

His offensive production has actually fallen about 21 points of wOBA below what Statcast would expect from the underlying events.

Crow-Armstrong provides a fascinating contrast.

\Delta \mathrm{wOBA}_{\mathrm{CrowArmstrong}} = 0.4013 - 0.3684 = 0.0329

His actual wOBA has exceeded his xwOBA by roughly 33 points.

That does not invalidate what he has accomplished. Runs that have already scored still count. It does suggest, however, that Alvarez’s underlying offensive performance is considerably stronger than the small difference between their total Offense numbers might initially imply.

James Wood might be the most frightening hitter behind Alvarez

Wood ranks third in total Offense at 33.68 runs, but that ranking almost understates what is happening.

His xwOBA is .4183, second only to Alvarez.

His barrel rate is 20.6 percent, the highest in the dataset. His hard-hit rate is 58.1 percent, also the highest.

Then there is his walk rate: 16.7 percent.

That is an unusual combination. Wood is not simply crushing baseballs. He is also refusing pitches frequently enough to force pitchers into the strike zone.

His main weakness remains strikeouts. His K% is 28.6 percent, which is high. Yet when he connects, the quality of contact is extraordinary.

The upper-right portion of Figure 5 is where we want to look. That region combines high barrel frequency with high expected offensive production.

Wood is there.

So is Alvarez, although Alvarez achieves his even higher xwOBA without matching Wood’s astonishing barrel percentage. Ohtani also occupies the elite region, while Pete Alonso combines tremendous hard contact with a .3946 xwOBA.

These are different offensive machines producing similar outcomes.

Bobby Witt Jr. is an important counterexample

Bobby Witt Jr. does not appear among the very top total offensive producers. His Offense value is 20.88, which ranks 21st in the dataset, and his wRC+ is 121.5.

Yet his xwOBA is .3820.

His actual wOBA is only .3488.

\Delta \mathrm{wOBA}_{\mathrm{Witt}} = 0.3488 - 0.3820 = -0.0332

That is one of the largest negative gaps in the entire sample.

Witt has also contributed 7.36 baserunning runs, third best among these hitters. His overall WAR is 6.31, second only to Crow-Armstrong in the dataset.

This is precisely why the definition of “best” matters.

If we ask who has produced the most offense, Witt is not at the very top. If we ask who has played the most valuable all-around baseball, the answer changes substantially. If we ask whose contact suggests better offensive results than he has received, Witt suddenly becomes extremely interesting.

One leaderboard cannot answer all three questions.

A standardized look at offensive styles

To compare very different statistics on a common scale, I standardized six measures across the full 142-player sample.

For player and metric :

z_{i,m} = \frac{ x_{i,m} - \overline{x}_m }{ s_m }

For strikeout rate, I reversed the sign so that positive values always represent the favorable direction:

z^{*}_{i,\mathrm{K\%}} = - z_{i,\mathrm{K\%}}

The resulting profiles are shown below.

The contrast is useful.

Alvarez is approximately 3.5 standard deviations above the sample mean in wRC+ and almost 3.8 standard deviations above average in xwOBA. There is no obvious weakness at the plate. His major negative component comes from baserunning.

Crow-Armstrong’s profile is more balanced between hitting and running. His ISO is exceptional, and his baserunning value is roughly two standard deviations above the sample mean, but his xwOBA is much less extreme than his actual offensive results.

Wood has perhaps the most intriguing shape of all. His xwOBA, power, and walk rate are spectacular, while his strikeout rate pulls strongly in the opposite direction.

Bryce Harper shows yet another profile. He combines elite plate discipline with strong expected production, but negative baserunning reduces his total contribution.

Four hitters survive almost every test

One way to avoid becoming overly dependent on a particular metric is simply to ask which players remain near the top regardless of how we look.

Only four hitters rank in the top 10 in Offense, wRC+, and xwOBA:

Player Offense Rank wRC+ Rank xwOBA Rank
Yordan Alvarez 1 1 1
James Wood 3 4 2
Shohei Ohtani 4 8 3
Bryce Harper 8 5 5

That is a formidable group.

But even within this smaller group, Alvarez separates himself. He is not merely first by one convenient statistic. He is first in total Offense, first in wRC+, first in wOBA, and first in xwOBA.

He also leads this dataset in WPA and RE24.

Different approaches keep returning the same name.

So who has been the best offensive player?

Through September 5, I think the answer is Yordan Alvarez, and the case is unusually strong.

Pete Crow-Armstrong has been close in actual total offensive value and has added considerably more on the bases. His season is extraordinary. But once expected production is introduced, the small gap between them begins to look larger. Alvarez’s .450 xwOBA suggests that the quality of his offensive performance may be even better than the already spectacular results indicate.

James Wood deserves special attention as well. If the question changes from “Who has produced the most?” to “Whose underlying offensive profile frightens me most going forward?”, Wood becomes a serious candidate. A 20.6 percent barrel rate, 58.1 percent hard-hit rate, .418 xwOBA, and 16.7 percent walk rate is an extraordinary collection of traits.

Ohtani and Harper complete what might reasonably be called the most robust elite group in the data.

The larger lesson, however, is about measurement itself. Statistics do not merely rank players. Different statistics ask different questions.

Offense tells us how much was produced. wRC+ tells us how efficiently it was produced. xwOBA tells us what the underlying contact suggests should have been produced. Baserunning shows us how value can be created after the ball leaves the bat.

When all of those measurements point toward the same player, the conclusion becomes difficult to avoid.

So far in 2026, Yordan Alvarez has been baseball’s most impressive offensive producer. That is a fact.

 

Stem-and-Leaf Plots and Histograms: Two Views of the Same Data

I downloaded 2026 MLB OPS data through the end of August to create the following figures. I always start with Stem-and-Leaf Plots when I begin a study. This short post is not about the distribution of OPS in Major League Baseball; I simply want to illustrate the difference between Stem-and-Leaf Plots and Histograms. Why? I think it is a useful exercise, especially for aspiring scientists or anyone curious about how to approach data.

A histogram and a stem-and-leaf plot can look remarkably similar. Both are designed to reveal the distribution of a numerical variable. Both can show where observations cluster, where the tails extend, and whether the distribution appears symmetric or skewed.

But they do not show the data in quite the same way.

Consider the OPS values for the 142 players in this dataset. The values range from .542 to 1.035, with a mean of .756 and a median of .747. When those values are placed in a histogram, the overall shape of the distribution becomes immediately visible.

The histogram excels at this.

Grouping OPS values into intervals provides a quick visual summary of where most players are concentrated. We can see the center of the distribution, its spread, and the relatively small number of players occupying the extreme upper and lower ends. If the primary question is “What does this distribution look like?”, the histogram is difficult to beat.

There is a cost, however. The individual observations disappear.

Suppose a histogram bar represents OPS values between .750 and .775. We know how many players fall within that interval, but we cannot see their actual OPS values. A player with a .751 OPS and another with a .774 OPS are simply members of the same bin.

The stem-and-leaf plot preserves that information.

Using a key such as

0.7 | 5 = .75 OPS

we can still see the shape of the distribution, but the leaves retain the underlying observations. A cluster of leaves tells us that many players occupy a particular region, while the individual digits allow us to reconstruct their approximate OPS values.

The split-stem version offers an especially useful compromise. Each tenth is divided into two rows, with leaves 0 through 4 on one row and 5 through 9 on the next. It creates more visual separation than the highly condensed stem-and-leaf plot, without producing the very long display that results from using hundredths as individual stems.

This highlights an important distinction between the two techniques.

A histogram emphasizes shape.

A stem-and-leaf plot emphasizes shape while preserving the data.

That makes the histogram particularly effective for larger datasets and quick visual comparisons. The stem-and-leaf plot is often more revealing with small or moderate datasets because it allows us to move back and forth between the distribution and the observations that created it.

Neither graph is inherently better.

They answer slightly different questions.

The histogram asks us to look at the forest. The stem-and-leaf plot lets us see the forest while still being able to identify many of the trees.

For exploratory data analysis, there is considerable value in looking at both.

POSTSCRIPT

I took a closer look at the two plots from above. The most interesting thing about them is the extreme outlier evident at the higher end. This is exactly what Exploratory Data Analysis is for. The outlier is Yordan Alvarez of the Houston Astros. Take a look at the Box Plot I created based on the same data as the Stem-and-Leaf Plot and Histogram.

The small circle on the right-hand side represents Alvarez. He is having an exceptional season; his OPS is a bit of an anomaly. His performance is far above what everyone else in the league has achieved. In this instance, the Box Plot is my favorite visualization.  It clearly shows how unusual Alvarez has been this year.

 

The Stem-and-Leaf Plot: Seeing the Distribution Without Losing the Data

Statistical graphics are often acts of compression.

A histogram takes individual observations and places them into intervals. A box plot goes further, reducing a distribution to a handful of landmarks. A density plot replaces the observations with a smooth estimate of their underlying shape. Each representation makes the data easier to understand by deliberately discarding some of their detail.

Usually, that is exactly what we want. A dataset containing ten thousand observations would be nearly useless if our only option were to stare at ten thousand numbers.

The stem-and-leaf plot takes a different approach.

It organizes the data visually while preserving the individual observations. Instead of choosing immediately between the raw numbers and a summarized picture of them, it gives us something of both. We can see the shape of the distribution, but we can also recover the values that created that shape.

That combination is unusual. It is also the reason the stem-and-leaf plot deserves more respect than its modest reputation might suggest.

The method is often introduced early in statistics courses and then quietly abandoned. Students learn to separate a number into a stem and a leaf, complete a few exercises, and move on to histograms, box plots, scatterplots, and regression. The stem-and-leaf display can therefore acquire the reputation of being a teaching device rather than a serious statistical tool.

That misses its deeper value.

A stem-and-leaf plot raises one of the most important questions in data analysis:

How much of the original data should we give up in order to see its structure more clearly?

That question extends far beyond stem-and-leaf plots. It lies near the heart of statistics itself.

Seeing the Distribution Without Losing the Observations

Consider a small dataset:

12, 14, 17, 21, 22, 22, 25, 28, 31, 34, 38, 39

Presented as a row of numbers, the observations are complete. Nothing has been lost.

But the structure is not particularly obvious.

We can improve matters immediately by sorting them:

12, 14, 17, 21, 22, 22, 25, 28, 31, 34, 38, 39

Now the range becomes easier to see. Repeated values are noticeable. The center is beginning to emerge.

A stem-and-leaf plot reorganizes the same information:

1 | 2 4 7

2 | 1 2 2 5 8

3 | 1 4 8 9

with the key:

2 | 5 = 25

The tens digit forms the stem. The ones digit forms the leaf.

The observation 25 becomes:

2 | 5

The observation 38 becomes:

3 | 8

And so on.

The important point is not merely that the data have been rearranged. It is that the rearrangement reveals a distribution.

The twenties contain the greatest concentration of observations. There are fewer values in the teens. The thirties contain several observations but are not as dense as the twenties. The values extend from 12 to 39.

We are seeing both the forest and the trees. That is difficult to accomplish with most statistical graphics.

Figure 1. From Raw Data to a Stem-and-Leaf Plot

Figure 1 illustrates the transformation from an unsorted list to an ordered list and finally to the stem-and-leaf display. The important thing to notice is that the observations survive every step. Their arrangement changes, but the values themselves remain available.

A histogram would already have begun to sacrifice that information, while a box plot would sacrifice much more.

The stem-and-leaf plot delays the sacrifice, and I have always found that noteworthy.

What the Plot Reveals

Once the observations are arranged in a stem-and-leaf display, several properties of the distribution become immediately visible.

Consider another example:

1 | 8 9

2 | 0 1 1 2 3 4 5 7 8 9

3 | 0 0 1 2 3 4 6 8

4 | 1 3

5 | 9

Most of the data fall in the twenties and thirties. The distribution becomes thinner as we move toward either extreme.

The value 59 stands apart. That does not prove that 59 is an outlier in any formal statistical sense. Nor does it tell us why the value is unusual. It may be a legitimate observation, a measurement error, a member of another population, or simply an improbable value from the same population.

But the display has done something useful. It has drawn our attention to it, and that is a central purpose of exploratory data analysis.

A good exploratory graphic does not necessarily answer the question at hand. Often, its most important contribution is telling us which question to ask next.

Center

Because the values are ordered, the median can be obtained directly.

Return to the twelve observations:

12, 14, 17, 21, 22, 22, 25, 28, 31, 34, 38, 39

There are twelve values, so the median is the average of the sixth and seventh observations:

\mathrm{Median} = \frac{22+25}{2} = 23.5

There is no need to reconstruct the dataset from a graph. The observations are already there.

Spread

The minimum and maximum are equally obvious.

\mathrm{Range} = x_{\max} - x_{\min}

For this dataset:

\mathrm{Range} = 39-12 = 27

Again, the calculation is simple because the display has preserved the values.

Quartiles

The same ordered structure gives us access to the quartiles. Once the first and third quartiles have been identified, the interquartile range follows:

\mathrm{IQR} = Q_3-Q_1

This establishes a direct connection between stem-and-leaf plots and box plots.

The box plot may look like an entirely different graphical object, but its landmarks come from the same ordered data. The stem-and-leaf display allows us to see the observations from which those landmarks were calculated.

The box plot presents the summary; the stem-and-leaf plot reveals the evidence behind it.

Figure 2. From Stem-and-Leaf Plot to Box Plot

Figure 2 makes that relationship explicit. The same dataset appears first as individual observations and then as a box plot. The first quartile, median, and third quartile have not magically appeared; they were extracted from the ordered values.

This is interesting and important because every summary statistic conceals the observations from which it was derived.

The stem-and-leaf plot keeps them visible a little longer, which can be very useful.

What Gets Lost When We Summarize

There is nothing inherently wrong with losing information; statistics would scarcely be possible without it.

The mean of a dataset compresses many observations into one number. The standard deviation compresses the pattern of variation into another. A regression model may reduce thousands of observations to a handful of coefficients. A box plot may represent a large distribution with a few lines and perhaps some points indicating unusual observations.

Compression makes patterns manageable, but it always comes at a cost. The stem-and-leaf plot is interesting because that cost is unusually small.

The histogram

Suppose we have the following observations:

21, 22, 22, 23, 24, 25, 27, 28, 29

If we construct a histogram using a single bin from 20 through 29, every observation appears inside the same bar we learn that there are nine observations in the twenties.

We no longer know where they are within the twenties.

Another dataset might be:

20, 20, 20, 20, 25, 29, 29, 29, 29

Under the same binning scheme, it could produce the same bar, yet the internal arrangements of the two datasets are dramatically different.

The histogram has not made an error; it has answered a different question.

It tells us how many observations fall within an interval; it does not guarantee that the observations themselves are preserved.

Figure 3. Same Histogram, Different Data

Figure 3 demonstrates the problem. Two datasets can produce identical bin counts even though the locations of the observations within those bins differ considerably.

The broader lesson is important. A histogram is partly a function of the data and partly a function of the bins we choose. Changing the bin width may change the apparent shape.

That does not make histograms unreliable. It simply means that their visual structure depends on the analyst’s decision.

The box plot

The box plot compresses even more aggressively.

Two datasets can have similar quartiles and medians while possessing noticeably different internal structures. One might be evenly distributed, another might contain clusters, a third could contain a large gap. Yet their box plots may look surprisingly similar.

Figure 4. Similar Box Plots, Different Internal Structure

The value of Figure 4 lies in the contrast. The box plots suggest similarity because the important quartile landmarks are alike. The stem-and-leaf displays reveal differences inside those landmarks.

Neither representation is wrong, they are simply preserving different information.

The box plot asks:

Where are the important summary locations of the distribution?

The stem-and-leaf plot asks:

How are the observations actually arranged?

Those are different statistical questions.

Gaps, Clusters, and Shape

One of the strongest uses of a stem-and-leaf display is its ability to reveal local structure.

Consider:

1 | 1 2 3 4

2 | 0 1 2 3

3 |

4 | 5 6 7 8

The thirties are empty, that absence is difficult to overlook.

Perhaps the data contain two populations. Perhaps some process separates low observations from high ones. Perhaps the apparent gap is nothing more than chance.

The display cannot tell us which explanation is correct; it can tell us that something interesting has happened.

A box plot may conceal the gap almost completely. A histogram may show it, but whether it does can depend strongly on the chosen bins.

The stem-and-leaf display shows the absence directly because the missing observations remain missing in plain sight.

Figure 5. A Distribution With a Gap

Figure 5 compares the same distribution using a stem-and-leaf plot, a histogram, and a box plot. Each display is useful, but the missing region is most literal in the stem-and-leaf representation.

This is one reason such plots are particularly valuable with small datasets. At that scale, individual observations still matter.

A gap consisting of five missing values may be statistically and scientifically interesting. When the dataset contains five million observations, preserving every value becomes much less useful. Scale changes the problem.

Resolution Is a Choice

Stem-and-leaf plots may appear objective because the observations themselves remain present. Yet even here, the analyst makes decisions.

Suppose many observations fall in the twenties:

2 | 0 0 1 1 2 2 3 4 4 5 5 6 6 7 7 8 8 9 9

The row is crowded.

We can split the stem:

2 | 0 0 1 1 2 2 3 4 4

2 | 5 5 6 6 7 7 8 8 9 9

The first line contains leaves from 0 through 4.

The second contains leaves from 5 through 9.

Nothing about the underlying data has changed. Only the visual resolution has changed.

Figure 6. Ordinary Stems Versus Split Stems

This is closely related to the choice of bin width in a histogram. Broad histogram bins conceal local variation. Narrow bins reveal more detail but can make the plot noisy.

The same tradeoff appears here. Large stems may hide structure while very fine stems may fragment the display.

There is no universal setting that is always best. This point is worth emphasizing because it applies to nearly every form of visualization. Graphs do not simply reveal data, they interpret them.

A histogram requires bins. A density plot requires a smoothing parameter. A map requires a projection. A box plot requires conventions about whiskers and outliers. Even axis limits can affect how dramatic a pattern appears.

The stem-and-leaf plot is no exception. Retaining the observations does not eliminate judgment; it merely makes one kind of information loss less severe.

The importance of the key

Another apparently trivial detail is actually essential.

Suppose we see:

12 | 3

What does it mean?

123?

12.3?

1.23?

The plot cannot tell us.

The key must.

For example:

12 | 3 = 12.3

or:

12 | 3 = 123

A stem-and-leaf display without a clear key is incomplete. This becomes especially important when the data contain decimals.

Suppose the observations are:

1.2, 1.4, 1.7, 2.1, 2.2, 2.8

We might display them as:

1 | 2 4 7

2 | 1 2 8

with the key:

1 | 2 = 1.2

The structure is unchanged. Only the decimal interpretation has moved.

Negative values can also be displayed, though the notation becomes less intuitive because the direction of numerical ordering must be handled carefully. In some situations, a dot plot or histogram will be easier to read. No statistical graphic is best in every circumstance.

Comparing Two Distributions

Stem-and-leaf plots are particularly interesting when comparing two groups. A back-to-back display places the stems in the center and the leaves of the two groups on opposite sides:

Group A Stem Group B

8 5 2 1 3 7

9 7 4 2 2 1 2 5 6 8

8 6 3 1 3 2 4 7 9

5 2 4 1 6 8

Now two distributions can be examined simultaneously without reducing either one to a few summary statistics. We can compare centers and spreads. We can look for differences in skewness, clusters, gaps, and extreme values. We can also see the degree of overlap.

Figure 7. Back-to-Back Stem-and-Leaf Plot

This makes the back-to-back display a useful companion to side-by-side box plots.

Suppose two box plots have different medians and relatively little overlap in their interquartile ranges. That may suggest an important difference between their central distributions.

The box plots tell us that the difference exists. The stem-and-leaf plots may help us understand what created it.

Perhaps one entire distribution has shifted upward. Perhaps the groups share a common lower range but differ in the upper tail. Perhaps one contains a cluster that the other lacks. Perhaps the apparent difference is being driven by only a few values.

The stem-and-leaf display does not replace the box plot; it provides another layer of evidence.

In exploratory work, different visualizations should often be treated as complementary rather than competitive.

When the Plot Stops Working

The great strength of the stem-and-leaf display eventually becomes its greatest weakness. It preserves the observations. That is wonderful when there are thirty observations. It may still be useful with a hundred. With a thousand, the display can become awkward; with a million, it becomes absurd.

A statistical graphic that insists on preserving every value cannot scale indefinitely.

Figure 8. The Scaling Problem

Figure 8 illustrates this transformation. With a small sample, individual leaves are meaningful. As the sample grows, the rows become crowded. Eventually, the attempt to retain every observation overwhelms the visual structure.

At that point, information loss becomes beneficial. The histogram succeeds precisely because it does not care about every observation. The box plot succeeds because it compresses even further. A density plot succeeds by replacing the observations with an estimate of shape.

An empirical cumulative distribution function can summarize the proportion of observations below each value without reproducing the data individually. The appropriate visualization changes with the scale of the problem.

This is an important principle. Plainly stated, more information is not always better. Sometimes the purpose of analysis is to decide which information can safely be ignored.

A Spectrum of Compression

We can think of several common representations as occupying positions along a continuum:

Raw Data

Stem-and-Leaf Plot

Histogram

Box Plot

This is not an absolute ranking. A histogram and box plot summarize different properties, and under some circumstances one may preserve something that the other does not.

Still, the general direction is useful.

As we move across the spectrum, individual detail decreases and compression increases.

Figure 9. The Compression Spectrum

Raw data preserve everything but may reveal very little. The stem-and-leaf plot organizes the observations while preserving their recoverability. The histogram sacrifices exact values to make the distribution’s shape more apparent. The box plot compresses the distribution into a compact set of summary landmarks that can be compared quickly across many groups.

Each step gains something, and each step loses something. The question is whether what we gain is more useful than what we give up.

That is a statistical judgment, not merely a graphical one.

Same Numbers, Different Stories

One of the dangers of statistical summaries is that they can create a false sense of completeness. Suppose two datasets have the same mean, that does not mean they have the same shape.

Suppose they have the same mean and standard deviation. They can still differ. They may contain different clusters, gaps, levels of symmetry, or tail behavior. Even similar quartiles do not guarantee similar distributions.

Figure 10. Same Mean and Spread, Different Shape

Figure 10 illustrates the idea using several small datasets with identical means and population standard deviations but noticeably different arrangements.

This is a recurring lesson in exploratory data analysis. A statistic is a description of the data, it is not the data.

The mean does not tell us everything, The standard deviation does not tell us everything.

The correlation coefficient does not tell us everything.

Even a sophisticated model does not contain every feature of the observations from which it was fitted. That is why visualization matters. And it is why looking at the data before summarizing them remains such an important habit.

Tukey and the Logic of Exploratory Data Analysis

The stem-and-leaf plot is closely associated with John Tukey and the tradition of exploratory data analysis. That connection is more important than the plot itself. Exploratory data analysis begins from a simple but powerful premise: before we impose a formal model on the data, we should examine what the data are trying to tell us.

Look for patterns. Look for exceptions, gaps, clusters, and asymmetry. Look for observations that refuse to behave like the rest.

This may sound obvious now, but it represents a distinctive way of thinking about statistics. Formal statistical analysis often begins with a model or hypothesis. Exploratory analysis begins with observation.

The stem-and-leaf plot fits that philosophy almost perfectly because it does not rush to compress the evidence. It organizes the values just enough for their structure to become visible.

The box plot, another graphic strongly associated with Tukey’s exploratory tradition, takes the next step. It compresses data. That is not a contradiction; it is a progression.

First we examine the observations, then we identify the structure. Then we decide which features can be summarized without losing what matters.

The sequence is important.

Statistics as the Art of Useful Loss

There is a larger lesson here. Statistics is often described as the science of learning from data. That is certainly true, but much of statistics might also be understood as the art of useful information loss.

A dataset may contain thousands or millions of values. We calculate a mean and most of the data disappear. We calculate a standard deviation and more structure is compressed.

We fit a regression line and thousands of individual points may become an intercept and a slope.

We construct a box plot and an entire distribution becomes a handful of positions on an axis.

Why would we do this?

Because raw information and useful information are not the same thing. If every observation were equally important at every stage of analysis, statistics would have little purpose. We could simply preserve the dataset forever and refuse to summarize it.

But human understanding requires structure. Compression makes structure possible. The challenge is deciding what can be discarded. Too little compression leaves us drowning in details. Too much compression can erase the phenomenon we are trying to understand.

That is why Figure 9, the compression spectrum, is more than a comparison of graphical techniques.

It represents a fundamental statistical tradeoff. The stem-and-leaf plot occupies an intriguing position near the beginning of that spectrum.

It says: Organize first. Discard later.

Why the Stem-and-Leaf Plot is Still Useful

It would be easy to dismiss the stem-and-leaf display as a relic from an earlier era of statistics.

Modern software can generate histograms, density plots, violin plots, ECDFs, box plots, and interactive visualizations almost instantly. Datasets are also vastly larger than the small samples for which stem-and-leaf displays are best suited.

Those observations are fair. The plot has limits; it possibly does not scale well.

Its notation can become awkward with complex measurements; it is rarely appropriate for extremely large datasets. There are many situations in which another graphic will be better.

Yet none of that makes the stem-and-leaf plot obsolete. Its greatest value may now be conceptual. It shows us what happens at the moment data become visualization.

We begin with individual observations. We reorganize them and structure appears. And remarkably little has been lost.

That makes the plot a useful bridge between raw data and statistical abstraction. It also teaches an important discipline. Do not summarize too quickly.

A mean may hide a gap. A standard deviation may hide clustering. A box plot may hide multimodality. A histogram may hide internal structure within its bins. A fitted model may hide observations that do not conform to its assumptions.

Before reducing the data, look at them. The stem-and-leaf plot makes that instruction almost literal.

The Forest and the Trees

There is a familiar warning about failing to see the forest for the trees. Statistics often presents the opposite danger. We may become so successful at seeing the forest that we forget the trees were ever there.

Summary statistics are powerful because they allow us to step back, while graphs allow us to recognize shape, and models allow us to identify relationships. All of these are essential.

But every abstraction moves us farther from the observations. The stem-and-leaf plot is unusual because it takes only a small step. It reveals the distribution without completely surrendering the values that created it.

For a modest dataset, that can be extraordinarily useful. We can see the center and the spread. We can see gaps, clusters, repeated values, and possible outliers. We can estimate quartiles and construct a box plot.

We can compare groups and if something surprises us, the original observations are still sitting there in front of us. That may be the plot’s most enduring lesson.

Statistics is not simply about reducing data.

It is about reducing data carefully. The goal is not to preserve everything forever, nor to summarize as aggressively as possible. The goal is to retain what matters long enough to understand what the data are saying.

The stem-and-leaf plot does that unusually well. It shows us the forest, and, for a little while longer, it lets us keep the trees.

 

The All-Star Who (Initially) Did Not Look Like One

The All-Star Who (Initially) Did Not Look Like One

Did Travis Bazzana Really Deserve His 2026 Selection?

When Travis Bazzana was named to the 2026 American League All-Star team, my first reaction was surprise. I mean, I was shocked. I didn’t expect him to make the team, did you?

Bazzana had certainly been good. But an All-Star already? He had not even reached the major leagues until April 28. By the time the All-Star roster was selected, he had played only 58 games and accumulated 249 plate appearances.

That made the selection worth examining. I really want to know how and why he made the team.

There was another reason to look closely. Bazzana was not simply added by Major League Baseball to make sure Cleveland had a representative. MLB’s official roster announcement specifically identified him as the American League’s second baseman selected through the Player Ballot, a vote involving players, managers, and coaches. Ernie Clement of Toronto had already been elected the starting second baseman by the fans.

So this was not merely a question of whether Bazzana was good. The more interesting question was whether the players themselves got it right.

Defining “Deserved”

There are several subtly different ways to ask whether someone deserved to be an All-Star. The first is simple: Was Bazzana playing at an All-Star level?

The second is more demanding: Was Bazzana one of the best American League second basemen at the time the roster was selected?

And the third is the hardest: Was he the most deserving reserve once Clement had already been elected the starter?

Those questions do not necessarily produce the same answer. A player can be performing at an All-Star level without being the single best choice available. I’ll take a look to see if he was the best choice.

For this analysis, I used the FanGraphs data available through July 4, immediately before the roster announcement. Clearly, using August statistics to judge a July decision would allow information that voters could not have possessed at the time to contaminate the analysis.

I also imposed a 150-plate-appearance minimum. That eliminated tiny samples while comfortably retaining Bazzana’s 249 plate appearances.

One player required special treatment. Ezequiel Duran did not appear in the FanGraphs second-base-filter export, but he was certainly an official AL second-base candidate on MLB’s ballot and appeared in the broader qualified American League FanGraphs exports. MLB’s June ballot updates even showed him among the leaders at the position. I therefore restored him to the comparison pool.

The First Surprise: Bazzana Was Good

The raw FanGraphs numbers immediately undermine the argument that Bazzana was a reputation-driven selection.

He was slashing:

.250/.341/.412

with 7 home runs, 12 stolen bases, an 11.6 percent walk rate, and a 20.9 percent strikeout rate.

More importantly, his 114 wRC+ indicates that his offensive production was roughly 14 percent better than league average after the adjustments incorporated into the metric. He had accumulated +6.15 offensive runs, +2.07 baserunning runs, and 1.43 WAR in only 249 plate appearances. Those are legitimate numbers.

They do not immediately prove that Bazzana should have been the reserve second baseman. But they establish something important at the outset. The selection was not absurd.

Table 1. Leading AL Second-Base Candidates Through July 4, 2026

Player PA wRC+ BsR Def WAR
Ezequiel Duran 297 107 +0.88 +8.00 2.22
Jazz Chisholm Jr. 334 98 +4.67 +4.83 2.10
Chase Meidroth 364 104 -0.29 +4.36 1.90
Kody Clemens 301 123 +0.68 -2.22 1.76
Travis Bazzana 249 114 +2.07 -0.91 1.43
Ernie Clement 342 108 -0.60 -1.08 1.38
Cole Young 357 107 -0.81 -1.37 1.37
Gleyber Torres 190 128 -2.91 -0.18 1.01

The FanGraphs exports place Bazzana behind Duran, Jazz Chisholm Jr., Chase Meidroth, and Kody Clemens in total WAR, but slightly ahead of Clement and Cole Young. Duran’s broader AL export shows 2.22 WAR, driven in considerable part by a very strong defensive rating.

Figure 1 changes the tone of the discussion. Figure 1 shows the FanGraphs WAR totals for the leading American League second-base candidates through July 4. The important point is not simply that Bazzana ranked fifth. He remained within the main cluster of serious candidates despite having considerably fewer plate appearances than most of the players ahead of him.

Bazzana was not the WAR leader. He was not even particularly close to Duran. Yet neither was he buried among mediocre players. He occupied the middle of a fairly compact cluster of serious candidates.

That is our first important result.

The Offensive Case for Bazzana

If we look only at offense, Bazzana’s case becomes considerably stronger.

Kody Clemens had the best combination of power and overall offensive production among the leading candidates, carrying a 123 wRC+ and .500 slugging percentage. Gleyber Torres produced an even higher 128 wRC+, although he had only 190 plate appearances and lost substantial value on the bases.

Bazzana sat immediately behind that offensive tier.

His .341 on-base percentage was particularly valuable. He was walking frequently, making enough contact, adding modest power, and creating additional value with his legs.

FanGraphs decomposes offensive value in a useful way:

\mathrm{Offense} = \mathrm{Batting} + \mathrm{Base\ Running}

For Bazzana:

\mathrm{Offense}_{\mathrm{Bazzana}} = 4.076 + 2.074 = 6.150

That is a strong total for someone with only 249 plate appearances. FanGraphs credited him with roughly 4.1 batting runs and another 2.1 runs on the bases.

The baserunning component should not be dismissed as decorative. Twelve stolen bases in 58 games certainly attracted attention, but BsR attempts to measure more than stolen bases alone. Bazzana was producing real value outside the batter’s box.

This is where his All-Star case begins to make intuitive sense. Players facing Cleveland were not necessarily thinking about Bazzana’s WAR ranking. They were seeing a hitter who controlled the strike zone, got on base, ran well, and was already producing above-average offense only weeks into his major-league career.

The Player Ballot may have been capturing something real.

Then Defense Changes Everything

The reason Bazzana’s total WAR was not higher was defense. FanGraphs credited him with -0.91 Def through the cutoff date. That figure combines fielding performance with positional value.

Conceptually:

\mathrm{Defense} = \mathrm{Fielding} + \mathrm{Positional}

For Bazzana, the underlying FanGraphs components were approximately -1.27 fielding runs and +0.36 positional runs, producing the -0.91 total. Compare that with Duran. Duran had only a 107 wRC+, lower than Bazzana’s 114. His total offensive value was +3.31 runs, also well below Bazzana’s +6.15.

But Duran’s FanGraphs defense was an extraordinary +8.00 runs. That defensive advantage propelled him to 2.22 WAR.

Jazz Chisholm Jr. followed a similar path. His 98 wRC+ was actually below league average, yet his baserunning and defense were excellent. FanGraphs gave him +4.67 BsR and +4.83 Def, enough to reach 2.10 WAR.

Kody Clemens was almost the mirror image. His offense was excellent, but FanGraphs rated his defense at -2.22 runs.

That produces an unusually interesting positional race.

Figure 2 plots offensive production against FanGraphs defensive value and shows just how different the candidates were. Bazzana sits on the offense-heavy side of the group, while Duran, Chisholm, and Meidroth received substantially more help from defense.

There was no single prototype for a valuable American League second baseman in early 2026.

Duran and Chisholm were being pushed upward by defense. Kody Clemens was being driven by his bat. Bazzana was getting most of his value from batting and baserunning while surrendering a relatively small amount defensively.

Figure 3 makes that contrast even clearer. The same contrast becomes more evident when the components are separated. Figure 3 compares batting, baserunning, and defensive value for the principal candidates. Bazzana’s profile is distinctive: strong batting value, meaningful baserunning value, and a relatively small defensive penalty.

This is also why defense must be included in the analysis. Ignoring it would artificially elevate Bazzana and Kody Clemens while punishing Duran, Chisholm, and Meidroth.

At the same time, there is a legitimate reason not to treat 80 games of defensive measurement as absolute truth. Defensive statistics stabilize slowly. Positioning, opportunity, scoring systems, and relatively small numbers of fielding plays can substantially affect estimates.

That creates a useful sensitivity test.

What Happens If We Trust Defense Less?

Instead of pretending that defensive measurement is perfectly precise, we can ask how the ranking changes as we gradually reduce its influence.

FanGraphs’ Runs Above Replacement framework can be represented approximately as:

\mathrm{RAR} = \mathrm{Offense} + \mathrm{Defense} + \mathrm{League} + \mathrm{Replacement}

WAR then converts those runs into wins:

\mathrm{WAR} = \frac{ \mathrm{RAR} }{ \mathrm{Runs\ per\ Win} }

For a sensitivity check, I recalculated the candidates while allowing defense to receive 0, 0.5, or its full FanGraphs weight. This is not a replacement version of WAR. It is simply a way to see how dependent the ranking is on defensive evaluation. Because defensive measurements are especially uncertain over partial seasons, it is worth asking whether the conclusion depends heavily on them. Figure 4 performs that sensitivity test by recalculating the candidates’ diagnostic value with defense given zero weight, half weight, and its full FanGraphs weight.

The result is revealing. When defense receives its full FanGraphs weight, Duran leads at 2.22 WAR, followed by Chisholm and Meidroth. Bazzana sits at 1.43.

When defense is given only half weight, Duran falls back toward the pack. Kody Clemens and Chisholm move near the top, while Bazzana remains around 1.48.

When defense is removed entirely, Kody Clemens becomes the leader and Bazzana rises to roughly third among the principal candidates.

Bazzana’s position is unusually stable. He does not need a favorable defensive estimate to make his case. Quite the opposite. Defense is holding his total down.

Duran’s candidacy is much more sensitive. His argument for being clearly superior to Bazzana rests heavily on trusting the defensive estimate.

That does not mean we should throw defense away. It means the gap between Duran and Bazzana is less certain than the raw 2.22 versus 1.43 WAR comparison initially suggests.

The Playing-Time Problem

There is another issue working against Bazzana. He simply had fewer opportunities.

Duran had 297 plate appearances. Chisholm had 334. Meidroth had 364. Clement had 342. Cole Young had 357.

Bazzana had 249.

That was not because he had been ineffective or injured for much of the major-league season. He did not make his debut until April 28.

One simple way to examine the effect of playing time is to normalize WAR to 600 plate appearances:

\mathrm{WAR}_{600} = 600 \left( \frac{ \mathrm{WAR} }{ \mathrm{PA} } \right)

For Bazzana:

\mathrm{WAR}_{600,\mathrm{Bazzana}} = 600 \left( \frac{ 1.432 }{ 249 } \right) \approx 3.45

That is not a projection. It does not mean Bazzana would necessarily have finished a full season with 3.45 WAR per 600 plate appearances.

It simply normalizes the rate at which value had accumulated. By that measure, Bazzana’s performance looks considerably more All-Star-like. His roughly 3.45 WAR-per-600 pace was competitive with the best players in the group.

The tradeoff is philosophical. Should an All-Star ballot reward the player who has accumulated the most value, or the player who has performed at the highest level while actually on the field?

There is no universally correct answer. If total contribution is the standard, Bazzana’s late arrival hurts him. If, on the other hand, quality of performance matters heavily, his case improves.

The Statcast Warning

There is, however, one piece of evidence that prevents me from becoming completely enthusiastic about Bazzana’s first-half numbers. His actual production was noticeably better than his Statcast expected production.

FanGraphs reported:

wOBA: .3319

xwOBA: .3025

SLG: .412

xSLG: .352

AVG: .250

xBA: .228

His barrel rate was only 4.2 percent, and his hard-hit rate was 36.1 percent.

We can express the difference between actual and expected wOBA as:

\Delta \mathrm{wOBA} = \mathrm{wOBA} - \mathrm{xwOBA}

For Bazzana:

\Delta \mathrm{wOBA}_{\mathrm{Bazzana}} = 0.3319 - 0.3025 = 0.0294

That is not trivial. It suggests that the quality of Bazzana’s batted-ball contact did not fully account for the results he produced.

But here again, context matters. Bazzana was not uniquely fortunate. A look at other players shows that Ernie Clement’s actual wOBA was about .324 while his xwOBA was only .272, an even larger gap. Duran also exceeded his xwOBA, although by a smaller amount, posting roughly .321 versus .303.

So Statcast weakens Bazzana’s case somewhat, but it does not destroy it. If anything, it warns us against treating the first-half offensive results of several second basemen as perfectly stable indicators of true talent. Figure 5 compares actual wOBA with Statcast expected wOBA for the second-base candidates. Players above the diagonal performed better than their contact quality would predict, while those below it underperformed their expected results. Bazzana is noticeably above the line.

The Duran Problem

At this point, one player becomes impossible to ignore, Ezequiel Duran.

If the question is simply, Who had accumulated the most FanGraphs value among the legitimate AL second-base candidates?, Duran has the cleanest answer.

He had 2.22 WAR, more plate appearances than Bazzana, and his offense was above average. And FanGraphs credited him with roughly eight defensive runs above average.
He was not an obscure candidate discovered only after the fact. MLB’s own All-Star voting updates showed Duran running near the top of the second-base voting while Bazzana was also among the leading candidates.

If I were selecting the reserve strictly from the FanGraphs value totals available at the time, I would choose Duran over Bazzana. But I would attach an asterisk to that conclusion. The size of Duran’s advantage depends heavily on defense, and half-season defensive values carry more uncertainty than half-season batting statistics. Take away the defensive separation, and Bazzana suddenly looks extremely competitive.

That makes the decision closer than the raw WAR totals imply.

And What About Clement?

There is also an interesting irony in the comparison with Ernie Clement. Clement was not merely elected to start. He became the leading American League vote-getter in Phase 1, automatically locking down the starting assignment at second base.

Statistically, however, Bazzana’s FanGraphs case was at least as strong.

Clement had 1.38 WAR to Bazzana’s 1.43. He had a 108 wRC+ to Bazzana’s 114. His baserunning contribution was negative while Bazzana’s was strongly positive. Both carried slightly negative FanGraphs defensive values in this snapshot.

Clement did have the playing-time advantage, 342 plate appearances compared with Bazzana’s 249. But if someone accepts Clement as a legitimate All-Star starter on performance grounds, it becomes difficult to argue that Bazzana was nowhere near All-Star caliber.

That comparison strengthens Bazzana’s defense considerably.

What Were the Players Seeing?

Statistics cannot tell us why players, managers, and coaches voted for Bazzana, but they allow us to construct a plausible explanation.

He was a rookie who had almost immediately become an above-average major-league hitter. He controlled the strike zone. He reached base. He ran aggressively and effectively. He was producing roughly 14 percent better offense than league average, and he had accumulated 1.4 WAR despite missing the opening month of the major-league season.

Players may also perceive certain skills more directly than statistical models do. They see bat speed, pitch recognition, baserunning pressure, defensive positioning, approach, and the quality of at-bats firsthand.

Reputation may have helped Bazzana. Being the first overall draft pick certainly made him familiar. But reputation alone is not a satisfactory explanation for the vote.

The numbers were already substantial enough to support it.

So, Did Travis Bazzana Deserve to Be an All-Star?

After reviewing the data, I would split the verdict into two parts. Did Travis Bazzana perform well enough to deserve serious All-Star consideration? Yes.

That conclusion is stronger than I expected when I began the analysis. A 114 wRC+, .341 OBP, positive baserunning value, and 1.43 WAR in only 249 plate appearances is an All-Star-caliber performance over the time he actually played.

His selection was not some inexplicable triumph of hype over production.

But the second question produces a different answer. Was Travis Bazzana clearly the most deserving reserve second baseman in the American League? No.

Ezequiel Duran had the strongest FanGraphs total-value case. Jazz Chisholm Jr., Chase Meidroth, and Kody Clemens also had legitimate statistical arguments. Bazzana’s relatively limited playing time and slightly negative defensive value keep him from being the obvious choice.

If I had been required to select one reserve strictly on the performance data available through July 4, I probably would have selected Duran.

Still, “probably Duran” is very different from “Bazzana did not deserve it.”

That distinction is the central finding of this exercise.

A Better Verdict

My original reaction to Bazzana’s selection was essentially: Really? Already?

The data changed that reaction. The best description of the selection is that it was not undeserved. Nor is it obvious but it is defensible.

Bazzana was already one of the strongest offensive second basemen in the American League. He added meaningful value on the bases. His WAR total was competitive despite substantially fewer plate appearances than most of the other leading candidates. Defense weakened his candidacy, and his Statcast-expected numbers suggest that some of his offensive production exceeded the quality of his contact.

Those are real reservations. But they are reservations about whether he was the best choice, not whether he belonged in the conversation.

That may be the most interesting thing about All-Star selections. The statistics often do not identify one incontrovertible winner. They reveal the tradeoffs.

Duran offered accumulated value and defense.

Chisholm offered defense and speed.

Kody Clemens offered offensive power.

Meidroth offered a balanced package with substantial defensive value.

Bazzana offered offense, on-base ability, speed, and an impressive rate of value accumulation in a shortened first half.

The Player Ballot chose Bazzana. The numbers say that choice was arguable yet reasonable.

My guess is that there will be many more All-Star games in Bazzana’s future. He is an impressive young player.

 

How Many Pitching Statistics Do We Really Need?

How Many Pitching Statistics Do We Really Need?

Correlation, Redundancy, and the Search for Independent Information in Baseball’s Pitching Metrics

Baseball has no shortage of pitching statistics.

ERA is still with us. So is WHIP. But they now share the stage with FIP, xFIP, SIERA, K-BB%, WAR, BABIP, ground-ball percentage, strikeout percentage, walk percentage, home-run rate, and an expanding collection of increasingly specialized measures.

That creates an interesting problem.

Are all these statistics actually telling us different things?

Or have we created many different ways of describing the same underlying pitching abilities?

The question became especially interesting after my study of WHIP. WHIP was very strongly associated with same-season ERA, but it was considerably weaker at predicting ERA one year later. FIP, xFIP, SIERA, and K-BB% all performed better as forward-looking measures.

That suggested a different question.

Perhaps we should stop asking which statistic is “best.”

Instead, we should ask:

How much independent information does each statistic actually contain?

That is the question I investigate here.

The Data

I used the same season-level FanGraphs dataset covering 2002 through 2025.

The data include season, innings pitched, ERA, FIP, xFIP, WAR, BABIP, home-run rate, and related pitching measures. A separate rate-statistics export includes WHIP, K%, BB%, K-BB%, ERA-, FIP-, xFIP-, FIP, xFIP, and SIERA. The batted-ball data add ground-ball percentage and related contact measures.

As in the WHIP study, I required at least 100 innings in a season:

IP_{i,y} \geq 100

That produced 1,699 qualifying pitcher-seasons.

For the predictive analysis, a pitcher had to reach 100 innings in consecutive seasons:

IP_{i,y} \geq 100 \quad \text{and} \quad IP_{i,y+1} \geq 100

That left 927 consecutive-season pairs involving 299 different pitchers.

Because no pitcher reached 100 innings during the shortened 2020 season, that year naturally drops out of the consecutive-season analysis.

The Metrics

The main study included:

ERA, WHIP, FIP, xFIP, SIERA, K%, BB%, K-BB%, HR/9, BABIP, GB%, and WAR.

Some of these variables are obviously related.

One relationship is exact:

\mathrm{K\!-\!BB\%} = K\% - BB\%

This simple equation illustrates the broader issue surprisingly well.

K%, BB%, and K-BB% may occupy three separate columns on a leaderboard, but they do not represent three independent pieces of information.

K-BB% is constructed directly from the other two.

The relationships among FIP, xFIP, SIERA, ERA, WHIP, and the underlying pitching rates are more complicated.

But the same basic problem remains.

Different names do not necessarily mean different information.

First Look: The Correlation Matrix

The natural place to begin is with Pearson correlation.

Rather than displaying the variables in an arbitrary order, I used hierarchical clustering to place statistics with similar correlation structures near one another.

Figure 1. Correlation Structure of Major Pitching Metrics

Several relationships immediately stand out.

Metrics Correlation
xFIP and SIERA 0.970
K% and K-BB% 0.939
FIP and WAR -0.885
FIP and xFIP 0.873
SIERA and K-BB% -0.858
FIP and SIERA 0.854
ERA and WHIP 0.811

The relationship between xFIP and SIERA is extraordinary.

Their correlation is approximately:

r_{\mathrm{xFIP},\mathrm{SIERA}} = 0.970

Squaring that correlation gives:

R^2 = (0.970)^2 \approx 0.941

So roughly 94 percent of their observed variation is shared in a simple linear sense.

That does not make xFIP and SIERA identical. They are calculated differently and emphasize somewhat different aspects of pitching.

But statistically, they move together to an extraordinary degree.

ERA and WHIP tell a similar, though less extreme, story. Their same-season correlation is approximately 0.81.

WHIP and ERA look like different statistics.

In practice, they often move together.

Correlation Is Only the Beginning

Pairwise correlation cannot tell us everything.

A statistic may have only moderate correlations with several individual variables while still being highly predictable from all of them collectively.

This matters enormously in regression analysis.

Suppose we try to predict one variable using all the others. If that variable can already be reconstructed very accurately, then adding it to a large regression may provide very little truly new information.

A standard way to investigate this is the Variance Inflation Factor:

\mathrm{VIF}_j = \frac{ 1 }{ 1-R_j^2 }

Here, R_j^2 measures how well predictor (j) can itself be predicted by the other predictors.

A VIF near 1 suggests relatively little redundancy.

As VIF increases, multicollinearity becomes more serious.

I excluded K-BB% from this particular calculation because K%, BB%, and K-BB% have an exact mathematical dependency.

The remaining results were striking.

Figure 2. Multicollinearity Among the Pitching Metrics

Metric VIF
FIP 109.0
WHIP 93.1
K% 47.5
HR/9 46.9
xFIP 44.1
BB% 32.1
BABIP 29.2
SIERA 28.2
ERA 5.8
GB% 3.9

These values are enormous.

But they should not be interpreted as evidence that FIP, WHIP, or SIERA are bad statistics.

That is not what VIF measures.

The result, instead, shows that combining all these variables into a single regression equation creates extreme redundancy.

Several predictors are trying to explain the same underlying variation.

That will matter shortly.

What Should We Predict?

A statistic can look extremely impressive when it is asked to explain something happening in the same season.

Prediction is harder.

So, as in the WHIP study, I used next-season ERA as the principal target.

The general idea is:

\widehat{\mathrm{ERA}}_{i,y+1} = \beta_0 + \sum_{j=1}^{p} \beta_j z_{i,y,j}

where the predictors come from year (y), while the target is ERA in year (y+1).

The predictors were standardized:

z_{i,y,j} = \frac{ x_{i,y,j} - \overline{x}_{j} }{ s_j }

Standardization allows coefficients from variables measured on very different numerical scales to be compared more sensibly.

More importantly, I did not simply fit the models and report their in-sample (R2).

I used leave-one-season-out cross-validation.

One season was withheld.

The model was trained using all the other seasons.

Then it had to predict the observations belonging to the season it had not seen.

The procedure was repeated across the available seasons.

That creates a much more demanding test.

Which Single Metric Predicts Best?

Before constructing complicated regression models, it makes sense to give each statistic a chance by itself.

Figure 3. Which Single Metric Best Predicts Next-Season ERA?

The results were:

Predictor in year (y) Predictive (R^2) for ERA in (y+1)
SIERA 0.206
FIP 0.199
xFIP 0.195
K-BB% 0.178
K% 0.176
WAR 0.145
ERA 0.116
WHIP 0.101
HR/9 0.059
GB% -0.009
BB% -0.009
BABIP -0.012

SIERA wins.

But only narrowly.

FIP and xFIP are very close behind it.

That is exactly what we might expect from the correlation matrix. If several statistics contain much of the same information, their predictive performance should often be similar.

K-BB% may be the most impressive result in the table.

It is remarkably simple:

\mathrm{K\!-\!BB\%} = K\% - BB\%

Yet its predictive (R2) reaches approximately 0.178.

That is not far behind FIP, xFIP, and SIERA.

BABIP performs particularly poorly. Its cross-validated (R2) is slightly negative.

A negative predictive (R2) does not mean that higher BABIP magically predicts lower ERA.

It means something simpler.

For this particular prediction problem, the fitted BABIP model performs slightly worse than simply predicting the average ERA.

What Happens If We Put Everything Into One Regression?

Here is where things become interesting.

I constructed a full ordinary least-squares regression containing:

ERA, WHIP, FIP, xFIP, SIERA, K%, BB%, HR/9, BABIP, and GB%.

K-BB% was omitted because including K%, BB%, and K-BB% together would introduce exact linear dependency.

The model then produced something strange.

The standardized coefficient for WHIP was approximately:

\beta_{\mathrm{WHIP}} \approx -0.389

Taken literally, that would imply that a higher WHIP predicts a lower future ERA, once the other statistics are held constant.

That is not a sensible baseball interpretation.

It is a multicollinearity problem.

Figure 4. OLS, Ridge, and LASSO Coefficients

Ordinary least squares tries to divide explanatory credit among variables that contain overlapping information.

That can make individual coefficients unstable.

One variable gets a large positive coefficient.

Another highly related variable gets a negative coefficient.

A small change in the sample can alter them again.

The overall model can still predict reasonably well.

The individual coefficients, however, become difficult to interpret.

This is exactly why simply adding every available baseball statistic to a regression is not necessarily a good idea.

More variables do not automatically produce more knowledge.

Sometimes they produce more confusion.

Ridge Regression

Ridge regression offers one solution.

Rather than allowing coefficients to become arbitrarily large, Ridge penalizes them:

\min_{\beta} \left[ \sum_{i=1}^{n} \left( y_i-\widehat{y}_i \right)^2 + \lambda \sum_{j=1}^{p} \beta_j^2 \right]

The tuning parameter (lambda) determines how strongly the coefficients are shrunk toward zero.

When the predictors contain large amounts of overlapping information, this can make the model considerably more stable.

That is exactly what happened.

Under ordinary least squares, the standardized WHIP coefficient was approximately -0.389.

Under Ridge regression, it became approximately:

\beta_{\mathrm{WHIP,Ridge}} \approx 0.013

Essentially zero.

That is an important result.

The model is not saying WHIP is useless.

It is saying:

Once all the other pitching information is already known, WHIP contributes very little additional information about next-season ERA.

That is a very different statement.

LASSO: Let the Model Throw Statistics Away

LASSO takes the regularization idea further.

Instead of penalizing squared coefficients, it penalizes their absolute values:

\min_{\beta} \left[ \sum_{i=1}^{n} \left( y_i-\widehat{y}_i \right)^2 + \lambda \sum_{j=1}^{p} \left| \beta_j \right| \right]

This has an interesting consequence.

LASSO can set some coefficients exactly to zero.

That turns our original question into an empirical experiment.

Give the model all these pitching statistics.

Then ask:

Which ones does it decide it does not need?

In the full-sample fit using the cross-validated penalty, LASSO retained four nonzero predictors:

FIP

SIERA

K%

HR/9

ERA went to zero.

WHIP went to zero.

xFIP went to zero.

BB%, BABIP, and GB% went to zero.

This does not mean those discarded statistics contain no useful baseball information.

It means that, for predicting next-season ERA after the retained variables were already available, their additional contribution was small enough that LASSO discarded them.

That is precisely the kind of redundancy we set out to investigate.

How Stable Was LASSO’s Decision?

A single LASSO fit is useful, but correlated predictors can substitute for one another.

So I also examined how often each statistic survived across the outer cross-validation folds.

Figure 5. Which Metrics Does LASSO Keep?

The approximate selection frequencies were:

Metric Selected
FIP 100%
SIERA 100%
K% 100%
HR/9 90%
ERA 48%
BB% 24%
xFIP 19%
WHIP 10%
BABIP 10%
GB% 5%

Three statistics survived every time:

FIP, SIERA, and K%.

HR/9 survived in approximately 90 percent of the folds.

WHIP survived only about 10 percent of the time.

That is particularly interesting after the previous WHIP study.

WHIP is very good at describing current run prevention.

Yet once the regression already knows FIP, SIERA, strikeout rate, home-run rate, and the other variables, WHIP rarely contains enough unique predictive information to survive LASSO.

That is not a contradiction.

It is the distinction between useful information and unique information.

Does the Giant Model Actually Predict Better?

Now we arrive at the question that matters most.

Perhaps all this redundancy does not matter if the large model predicts much better.

Does it?

Figure 6. More Metrics Help, but Only a Little

The cross-validated results were:

Model Predictive (R^2)
ERA only 0.116
SIERA only 0.206
Five-skill model 0.215
All metrics, OLS 0.213
All metrics, Ridge 0.222
All metrics, LASSO 0.217

The five-skill model used K%, BB%, HR/9, BABIP, and GB%.

The best model was the full Ridge regression.

Its predictive (R2) was approximately:

R^2_{\mathrm{Ridge}} = 0.222

SIERA alone produced:

R^2_{\mathrm{SIERA}} = 0.206

The improvement was therefore:

\Delta R^2 = 0.222 - 0.206 = 0.016

Just 1.6 percentage points.

The RMSE tells essentially the same story.

SIERA alone produced an RMSE of about 0.737 ERA runs.

The full Ridge model reduced that to approximately 0.729.

The larger model is better.

But only a little.

That may be one of the most important results in the study.

We gave the model a large collection of modern pitching statistics.

Most of the additional information barely moved the prediction.

Once We Know SIERA, What Else Helps?

This provides another way to look at redundancy.

Start with SIERA, the strongest individual predictor.

Then add other statistics one at a time.

Figure 7. Once We Know SIERA, Most Extra Metrics Add Very Little

SIERA alone:

R^2 = 0.206

Add FIP and the result improves to approximately:

R^2 = 0.217

That is a genuine, although modest, improvement.

Add xFIP to SIERA and predictive performance actually slips slightly, to approximately 0.204.

Add WHIP and it falls to roughly 0.203.

Adding K-BB% changes almost nothing.

Even combining SIERA, FIP, and xFIP reaches only about 0.216.

This is a remarkably clean demonstration of statistical redundancy.

Three statistics are not necessarily three times as informative as one.

Sometimes the second statistic is largely repeating the first.

The third repeats them both.

Why xFIP and SIERA Are a Good Example

Consider again:

r_{\mathrm{xFIP},\mathrm{SIERA}} = 0.970

If two statistics move almost perfectly together, there is simply not much room for one to add entirely new predictive information after the other is already known.

They can still differ conceptually.

Their formulas can still have different purposes.

Their disagreements can even be analytically useful.

But conceptual difference does not guarantee statistical independence.

That distinction is central to this study.

What About WAR?

ERA is only one possible definition of future pitching success.

So I repeated the regression analysis using next-season WAR as the target.

Current WAR itself was the strongest individual predictor:

R^2_{\mathrm{WAR}_y,\mathrm{WAR}_{y+1}} \approx 0.336

SIERA reached approximately 0.252.

FIP reached about 0.243.

xFIP was around 0.235.

WHIP was considerably weaker at approximately 0.125.

Figure 8. Predicting Next-Season WAR

The multivariable models produced:

Model Predictive (R2)
WAR only 0.336
WAR + IP 0.336
Skill model + IP 0.297
All metrics OLS 0.373
All metrics Ridge 0.375
All metrics LASSO 0.377

Here the large models add somewhat more information.

The best LASSO model improves predictive (R2) from roughly 0.336 using current WAR alone to approximately 0.377.

That is a more noticeable gain than we observed when predicting ERA.

Still, the central lesson remains.

Adding many statistics helps.

But the improvement is nowhere near proportional to the number of variables added.

The Strange WHIP Coefficient Revisited

The negative WHIP coefficient in the ordinary regression is worth returning to because it demonstrates an important statistical point.

WHIP by itself predicts future ERA in the expected direction.

Higher WHIP is associated with higher future ERA.

But once we tell an ordinary regression to hold ERA, FIP, xFIP, SIERA, K%, BB%, HR/9, BABIP, and GB% constant, the WHIP coefficient becomes negative.

Those are two very different questions.

The simple regression asks:

What happens to future ERA when WHIP changes?

The giant multiple regression asks:

What happens when WHIP changes while an enormous collection of closely related pitching statistics somehow remains fixed?

That second scenario may have very little resemblance to an actual pitcher.

The variables are too interconnected.

OLS nevertheless tries to divide the shared information among them.

The resulting coefficient is mathematically legitimate.

Its baseball interpretation is questionable.

Ridge responds by shrinking the WHIP coefficient almost to zero.

LASSO simply removes WHIP.

Both approaches give us a more sensible picture of what the data are saying.

How Many Statistics Do We Really Need?

There is no universal answer.

It depends on the question.

If I want to know how successfully a pitcher kept runners off base, WHIP is excellent.

If I want to know how many earned runs he actually allowed, ERA tells me exactly that.

If I want a simple forward-looking estimate of next-season ERA, SIERA performed best among the individual metrics examined here.

If I want to squeeze out every additional bit of predictive accuracy, a regularized multivariable model performs somewhat better.

But the key word is somewhat.

Going from SIERA alone to a ten-variable Ridge regression improved predictive (R2) from about 0.206 to 0.222.

The sophisticated model wins.

Barely.

The Bigger Lesson

A modern baseball leaderboard can create an illusion of enormous amounts of independent information.

Twenty columns look like twenty facts.

Statistically, that may not be true.

Strikeout ability appears directly in K%.

It appears again in K-BB%.

It also enters FIP, xFIP, and SIERA.

Walks do the same.

Home runs affect FIP and other estimators.

BABIP influences hit-based measures.

ERA and WHIP share the consequences of allowing baserunners.

The columns multiply faster than the underlying baseball phenomena.

That is not a criticism of advanced statistics.

Different metrics were created to answer different questions. They emphasize different aspects of performance. They make different assumptions. They may be useful in different contexts.

The mistake would be assuming that because two statistics have different names, they must contain completely different information.

They often do not.

Conclusion

The original question was simple:

How many pitching statistics do we really need?

The answer is more interesting than I expected.

Many pitching metrics are strongly correlated.

Some are extraordinarily correlated.

xFIP and SIERA correlate at approximately 0.97. ERA and WHIP correlate at about 0.81. FIP, xFIP, SIERA, strikeout measures, walk measures, and home-run measures overlap so heavily that placing all of them in the same ordinary regression produces severe multicollinearity.

The VIF analysis shows the problem.

The unstable OLS coefficients make it visible.

Ridge regression controls it.

LASSO begins throwing redundant statistics away.

And the predictive results show just how little we lose by simplifying.

SIERA alone explains about 20.6 percent of next-season ERA variation under leave-one-season-out prediction.

A full ten-variable Ridge model improves that to approximately 22.2 percent.

Ten statistics are better than one.

But not by much.

That may be the most important conclusion.

The goal should not be to collect the largest possible number of metrics.

It should be to identify statistics that represent genuinely different dimensions of pitching.

After that point, we increasingly begin measuring the same underlying abilities again.

And again.

Just under different names.

 

6174: The Attractor Hiding in Four-Digit Arithmetic

Some numbers are famous because they are enormous. Others are famous because they appear everywhere.

6174 is famous because it is almost impossible to escape.

Take the number 3524.

Arrange its digits from largest to smallest, then from smallest to largest, and subtract:

5432 - 2345 = 3087

Now do it again.

8730 - 0378 = 8352

And again.

8532 - 2358 = 6174

We have arrived.

But the strange part is not that this particular sequence produced 6174. The strange part is that almost every four-digit number does.

And once 6174 appears, it refuses to leave:

7641 - 1467 = 6174

Do it again.

Again.

The number has become a mathematical destination.

Kaprekar’s Routine

The process is named after Dattatreya Ramchandra Kaprekar, an Indian schoolteacher and recreational mathematician who is generally credited with discovering the phenomenon in 1949. His work included several other unusual classes of numbers, but 6174 became his best-known discovery.

The procedure is wonderfully simple:

  1. Choose a four-digit number containing at least two different digits.
  2. Rearrange its digits into descending order.
  3. Rearrange those same digits into ascending order.
  4. Subtract the smaller number from the larger number.
  5. Treat the result as a four-digit number, adding leading zeros when necessary, and repeat.

We can describe the operation as a function. Let (D(n)) be the number formed by arranging the digits of (n) in descending order, and let (A(n)) be the number formed by arranging them in ascending order.

Then:

K(n) = D(n) - A(n)

The function (K) is the Kaprekar operation.

For 3524,

K(3524) = 5432 - 2345 = 3087

and repeated application gives

K^{3}(3524) = 6174

where (K3) means applying the Kaprekar operation three times.

That notation makes something important visible.

This is not merely a number trick.

It is an iterative system.

Every Road Leads to 6174

For four-digit decimal numbers containing at least two distinct digits, repeated application of the Kaprekar routine reaches 6174 in no more than seven iterations.

Seven is not just an upper bound that never actually occurs.

Consider 1004:

4100 - 0014 = 4086 8640 - 0468 = 8172 8721 - 1278 = 7443 7443 - 3447 = 3996 9963 - 3699 = 6264 6642 - 2466 = 4176

and finally,

7641 - 1467 = 6174

Exactly seven iterations.

Then the system stops changing.

Or, more precisely, it continues operating but remains at the same value:

K(6174) = 6174

In the language of dynamical systems, 6174 is a fixed point.

And that makes the number considerably more interesting than it first appears.

There Is One Important Exception

The rule requires at least two different digits.

Start with 4444, for example:

4444 - 4444 = 0000

And then:

0000 - 0000 = 0000

The same thing happens with 1111, 2222, 3333, and the other repeated-digit numbers.

They fall into 0000 instead of 6174.

So the claim is not that literally every possible four-digit string reaches 6174. It is that every four-digit decimal number with at least two distinct digits does.

That small qualification matters.

It also gives us another fixed point:

K(0000) = 0000

But 0000 is the trivial one.

6174 is where the interesting mathematics lives.

Something Is Being Destroyed

Why should thousands of apparently different numbers all collapse toward the same result?

The answer begins with the sorting operation.

Suppose the four digits, after sorting, are

a \geq b \geq c \geq d

The descending number is

1000a + 100b + 10c + d

while the ascending number is

1000d + 100c + 10b + a

Subtract them:

\begin{aligned} K(n) &= (1000a + 100b + 10c + d) \\ &\quad - (1000d + 100c + 10b + a) \end{aligned}

Collecting terms gives

K(n) = 999(a-d) + 90(b-c)

That equation is the part of the story I find most revealing.

The original number had four digits.

But after sorting and subtracting, the next value depends only on two differences:

a-d

and

b-c

Much of the information contained in the original number has vanished.

3524, 4253, 2435, 5342, and every other permutation of those same four digits are no longer different as far as the Kaprekar routine is concerned. Sorting makes them identical before the subtraction even begins.

The system is compressing its state space.

Repeatedly.

That begins to explain the convergence.

And Everything Becomes Divisible by 9

The same equation reveals another feature:

K(n) = 999(a-d) + 90(b-c)

Both 999 and 90 are divisible by 9.

Therefore,

K(n) \equiv 0 \pmod{9}

After only one iteration, every number produced by the routine is divisible by 9.

That is not a coincidence.

The descending and ascending numbers contain exactly the same digits, so they have the same digit sum. Two integers with the same digit sum have the same remainder modulo 9. Their difference must therefore be divisible by 9.

And sure enough,

6+1+7+4=18

so

6174 \equiv 0 \pmod{9}

Again, the routine is reducing the possibilities.

The apparent freedom of the starting number disappears very quickly.

Why 6174 Stays Put

Now look at 6174 itself.

Its digits in descending order are

7641

and in ascending order,

1467

so

7641 - 1467 = 6174

Using our difference formula gives the same result.

Here,

a=7,\quad b=6,\quad c=4,\quad d=1

Therefore,

a-d=6

and

b-c=2

so

\begin{aligned} K(6174) &= 999(6)+90(2) \ &= 5994+180 \ &= 6174 \end{aligned}

The number reproduces itself.

That is the defining feature of the Kaprekar constant.

How Quickly Do Numbers Fall Into 6174?

I decided to check every ordinary four-digit starting value from 1000 through 9999, excluding the nine repeated-digit numbers 1111, 2222, and so forth.

That leaves 8,991 starting values.

The convergence is remarkably fast.

Figure 1. Number of Kaprekar iterations required to reach 6174 for every starting value from 1000 through 9999 except the nine repeated-digit numbers. The value 6174 itself requires zero iterations.

The most common journey takes only three iterations. In my enumeration, 2,124 starting values reached 6174 in exactly three steps.

At the opposite end, 1,980 numbers required the full seven iterations.

Across all 8,991 valid starting values, the average iteration was approximately \overline{N}_{\mathrm{steps}} \approx 4.679

So although seven steps can be necessary, a randomly selected ordinary four-digit starting number is typically swallowed by the Kaprekar process considerably sooner.

The funnel is steep.

A Tiny Dynamical System

This is where 6174 stops being a curiosity about subtraction and becomes something more interesting.

Imagine every four-digit number as a node in a network.

Draw an arrow from each number to the number produced by one Kaprekar operation:

n \longrightarrow K(n)

Then draw another arrow:

K(n) \longrightarrow K^{2}(n)

and keep going.

Thousands of different starting points begin feeding into shared intermediate states. Those states merge into still fewer states.

Eventually the paths converge.

For almost the entire nontrivial four-digit system, they terminate at the same node:

6174

And that node points back to itself.

In dynamical-systems terminology, we can think of 6174 as an attractor, and the collection of starting values that eventually reach it as its basin of attraction.

Of course, this is a finite deterministic system rather than a continuous physical one. Nothing mysterious is pulling the numbers toward 6174.

The rules do it.

But that may be the most interesting part.

Order Emerging From an Almost Ridiculously Simple Rule

There is no probability in the Kaprekar routine.

No optimization.

No hidden choice.

No intelligence.

Take the digits. Sort them. Subtract. Repeat.

Yet a strong global pattern emerges.

That makes 6174 a miniature example of something that appears throughout mathematics and science: simple local rules can create surprisingly rigid large-scale behavior.

We see versions of this idea in cellular automata.

In iterative maps.

In fractals.

In differential equations.

In computer simulations.

The rules themselves can be almost embarrassingly simple. What happens after the rules are repeated may not be.

6174 gives us that lesson with arithmetic a child can perform.

6174 Is Also About Information

There is another way to look at the routine.

It destroys information.

Suppose I tell you that the output of one Kaprekar iteration is 3087.

Can you reconstruct the original starting number uniquely?

No.

Many different numbers can lead to the same result.

The map is many-to-one.

Once those different histories merge, the routine no longer remembers where they came from.

Iteration causes more and more paths to merge until the system has effectively forgotten almost everything about its initial state.

What survives is structure.

Eventually, for the four-digit decimal case, that structure is 6174.

Seen this way, Kaprekar’s routine is almost an information funnel.

Many possible states enter.

Far fewer distinct states survive.

Finally, almost everything exits through the same point.

It Is Not Just About Base 10

The phenomenon depends on the number of digits and the numerical base being used.

In ordinary base-10 arithmetic, there is a famous three-digit counterpart:

495

because

954 - 459 = 495

and the corresponding three-digit Kaprekar routine converges to 495 under the appropriate nontrivial starting conditions.

But change the number of digits or change the base, and the behavior can become much more complicated. Instead of one fixed point, systems may develop multiple fixed points or cycles.

In fact, Kaprekar dynamics remain an active mathematical subject. Recent work has studied the four-digit routine in other bases and found highly structured families of terminal cycles rather than simply reproducing the base-10 behavior of 6174.

So our little arithmetic trick opens the door to a much broader question:

What happens when a simple deterministic transformation is repeatedly applied to a finite universe of states?

That is no longer recreational arithmetic.

That is dynamics.

Why I Like 6174

6174 does not help us calculate the orbit of Mars.

It does not secure internet traffic.

It does not predict the stock market.

As far as I know, civilization would continue more or less unchanged if nobody had ever discovered it.

And yet I think that is part of its charm.

Kaprekar looked at ordinary decimal digits and asked what would happen if he performed a ridiculously simple operation again and again.

Most people would probably try it a few times, notice the pattern, and move on.

He paid attention.

There was structure hiding there.

That is one of the recurring pleasures of mathematics. The interesting thing does not always announce itself with an enormous theorem or an impossibly complicated equation.

Sometimes it is sitting inside four digits.

Sort them.

Subtract.

Repeat.

And no matter where you thought you were going, you discover that the road was leading to the same place all along.