The Men in the Middle: Who Was the Best Shortstop of the 1970s?

The shortstop position in the 1970s was very different from what we know today. Power was welcome, of course, but it was hardly the defining requirement. Teams generally expected their shortstops to catch the ball, make the throw, turn the double play, move runners when necessary, and contribute enough offensively to keep the lineup moving.

That makes the decade particularly interesting to study. There was no Cal Ripken Jr. yet, no Alex Rodriguez, no Derek Jeter, and certainly no modern collection of 25-home-run shortstops scattered throughout the league. The position still leaned heavily toward defense, speed, contact, and durability.

But who was actually the best?

The familiar names arrive quickly. Dave Concepción was the shortstop of the Big Red Machine. Mark Belanger was one of the most celebrated defensive players of his generation. Bert Campaneris brought speed and range to the Oakland dynasty. Larry Bowa became synonymous with Philadelphia baseball. Chris Speier, Don Kessinger, Freddie Patek, Rick Burleson, Toby Harrah, and others all made substantial claims of their own.

Rather than begin with reputation, I wanted to begin with the data. The result was clear at the top. The more interesting story, however, was what happened immediately behind the winner.

Defining a 1970s Shortstop

The study covers the seasons from 1970 through 1979. Players were identified by their appearances at shortstop, and the primary requirement was at least three qualifying shortstop seasons during the decade.

That qualification matters. A player who happened to spend one excellent season at shortstop should not necessarily outrank someone who occupied the position for most of the decade. At the same time, requiring eight or ten seasons would unfairly eliminate players whose careers began midway through the 1970s.

The three-season minimum produced a pool of 34 shortstops. As a sensitivity check, I later tightened the requirement to five qualifying seasons. The top five did not change, which matters when we consider how robust the final ranking really is.

The offensive model follows the same peer-adjusted Model C framework I have used in earlier positional studies. Each shortstop is compared with other shortstops from his own season rather than with hitters from another position or another era. The underlying model incorporates on-base percentage, isolated power, walk rate, strikeout avoidance, baserunning, runs scored, and RBI production.

For WordPress, the equation is:

\begin{aligned} \mathrm{OffensiveScore} &= z_{\mathrm{OBP}} + z_{\mathrm{ISO}} + z_{\mathrm{BB/PA}} + z_{\mathrm{LowSO/PA}} \\ &\quad + z_{\mathrm{NetSB/PA}} + z_{\mathrm{R/PA}} + z_{\mathrm{RBI/PA}} \end{aligned}

This approach changes the question in an important way. It does not ask whether Freddie Patek hit like a first baseman. He obviously did not. It asks how much offensive value Patek produced relative to the other men being asked to play shortstop at the same time.

Defense presents a more difficult problem.

The Lahman database gives us traditional fielding information, including assists, putouts, double plays, errors, games, and fielding percentage. It cannot provide the spatial information available in modern systems such as Outs Above Average, nor can it fully separate a shortstop from the pitchers, second basemen, field conditions, positioning, and scoring practices around him.

Still, we can construct a useful traditional defensive index:

\mathrm{TraditionalDefensiveScore} = z_{\mathrm{A/G}} + z_{\mathrm{PO/G}} + z_{\mathrm{DP/G}} + z_{\mathrm{FPct}} + z_{\mathrm{LowE/G}}

This is the same general traditional-defense construction used in our earlier positional work. It should be interpreted as a measure of performance in the defensive statistics available to us, not as a retroactive version of modern defensive runs saved.

Once the offensive and defensive dimensions were calculated, each was standardized and combined:

\mathrm{TwoWayScore}_{i} = z_{\mathrm{Offense},i} + z_{\mathrm{Defense},i}

That is the key to the study. We are not looking for the best hitter who happened to stand at shortstop, nor simply for the slickest glove. We are looking for the player who best combined the position’s two demands.

The First Answer

The ranking produces an unusually satisfying No. 1. Dave Concepción finishes first.

He is not first because the algorithm happens to reward longevity. He is also not first because of one extraordinary short peak. Concepción ranks first in overall two-way performance and first when the calculation is restricted to the best three-season peak. Offensively, he ranks second among the qualifying shortstops, while his traditional defense ranks fifth.

Figure 1. Best Shortstops of the 1970s: Two-Way Career Ranking.

Figure 1 shows how much separation Concepción creates. His standardized two-way career score of 3.36 comfortably exceeds Freddie Patek at 2.70 and Chris Speier at 2.55. Larry Bowa follows at 2.19, with Bert Campaneris rounding out the top five at 1.68.

The complete top ten is revealing:

Rank Player Offensive Rank Defensive Rank Peak Rank
1 Dave Concepción 2 5 1
2 Freddie Patek 3 7 2
3 Chris Speier 4 6 3
4 Larry Bowa 16 1 7
5 Bert Campaneris 5 12 8
6 Roy Smalley 6 13 4
7 Toby Harrah 1 27 6
8 Mark Belanger 25 2 9
9 Rick Burleson 9 9 5
10 Don Kessinger 7 11 11

The list immediately tells us something important. There was not one model of a successful 1970s shortstop. There were several.

Concepción simply inhabited the most valuable part of that landscape.

Concepción and the Power of Balance

Concepción’s offensive score for the decade is 31.83, second among the group, while his defensive score is 19.88, fifth. His strength is therefore not hidden inside a single category. He rates very well in both.

That balance is visible when offense and defense are plotted separately.

Figure 2. Different Paths to Shortstop Value in the 1970s.

The horizontal dimension in Figure 2 represents peer-adjusted offense. The vertical dimension represents traditional defense. Players in the upper-right portion of the chart are the rare shortstops who contributed substantially in both areas.

Concepción occupies exactly the region we would hope the best shortstop would occupy. Chris Speier and Freddie Patek are nearby, though neither reaches quite as far offensively. The graph makes Concepción’s advantage easier to understand than a ranking alone can.

It also reveals something much more interesting about the rest of the decade.

Larry Bowa and Mark Belanger live high on the defensive axis. Toby Harrah sits far to the right offensively but dramatically lower defensively. Freddie Patek, Chris Speier, Bert Campaneris, and Roy Smalley occupy a more balanced middle ground.

Different players reached value by very different routes.

The Freddie Patek Problem

If anything in this study surprised me, it wasn’t Concepción finishing first. It is Freddie Patek finishing second.

Patek ranks third offensively, seventh defensively, second in three-year peak, and second overall. His two-way career score of 2.70 places him ahead of Speier, Bowa, Campaneris, and Belanger.

That doesn’t look obvious if we start with conventional batting statistics. Patek was not a slugger. He did not put up the sort of raw offensive totals that jump from a baseball card.

But that is exactly why positional and seasonal context matter. Patek was being compared with other 1970s shortstops, not with Reggie Jackson or Willie Stargell. His ability to get on base, create runs with his legs, avoid being a complete offensive liability, and still provide useful defense matters more when evaluated against his positional peers.

The model does not claim that Patek was one of the best hitters in baseball. It says something narrower and more interesting: relative to what teams were receiving from shortstop, Freddie Patek created a surprisingly large amount of value.

That distinction is essential.

Chris Speier: The Quietly Complete Shortstop

Patek may be the surprise, but Chris Speier might be the cleanest example of an underrated two-way player.

Speier ranks fourth offensively and sixth defensively. Those rankings combine to place him third overall, and he is also third in three-season peak.

There is nothing extreme about his statistical profile. That is precisely what makes him interesting.

He does not need Harrah’s offensive advantage. He does not need Belanger’s defensive advantage. Speier remains near the top because there is no obvious weakness in his profile relative to other shortstops of the decade.

Baseball history sometimes remembers specialists more vividly than balanced players. Extraordinary power is memorable. Extraordinary defense produces highlight plays and reputations. A player who is simply very good at everything can be easier to overlook.

Speier appears to have been that kind of shortstop.

Larry Bowa and Mark Belanger

Now the geometry changes. Larry Bowa ranks only 16th offensively, but he finishes first defensively in the traditional model. The strength of that defense is sufficient to elevate him to fourth overall.

Mark Belanger pushes the idea even further. His offense ranks 25th, near the bottom of the serious candidates, but his defense ranks second. The combination still carries him all the way to eighth overall.

This is where the external evidence becomes particularly useful. Belanger won seven Gold Gloves during the decade. The Gold Gloves are not part of the model. They were added only after the rankings were calculated, so the agreement between the traditional defensive score and contemporary recognition serves as a useful independent check.

It does not prove that our defensive score measures everything correctly. Nothing built from traditional fielding statistics could. But if a model designed to detect exceptional traditional defense places Mark Belanger near the top, and contemporary observers repeatedly reached the same conclusion, that is reassuring.

Toby Harrah and the Other Extreme

If Bowa and Belanger demonstrate how defense can carry a shortstop, Toby Harrah demonstrates the opposite.

Harrah ranks first offensively among the qualifying shortstops of the decade. Defensively, however, he falls to 27th. That combination leaves him seventh overall, despite having the strongest offensive score in the study.

That makes Harrah one of the most informative players in the analysis. If we ranked shortstops by offense alone, Harrah would win. If we ranked them by traditional defense alone, Bowa would win. If we ask who combined the two dimensions most successfully, Concepción wins.

Those are three different questions. They deserve three different answers.

Peak Value Changes Less Than Expected

Career rankings can sometimes reward players who were merely good for a long time. A shorter peak measurement helps expose that problem.

For this study, I calculated each player’s best three-season two-way peak. The result strengthens rather than weakens Concepción’s case.

Figure 3. Career Value Versus Peak Value.

Concepción is again first. Patek is second and Speier third, giving us the same top three produced by the decade-long measure.

Roy Smalley becomes especially interesting here. He ranks sixth in decade value but fourth in peak value, despite having only four qualifying offensive seasons in the 1970s. Rick Burleson similarly rises to fifth in peak performance while ranking ninth over the full decade.

That distinction is useful because a decade boundary is artificial. Smalley did not organize his career around our decision to begin counting on January 1, 1970, and stop after the 1979 season. Younger players entering late in the decade naturally have less opportunity to accumulate career value.

Peak analysis gives them a second way into the conversation.

The Clusters of Shortstop Greatness

A ranking forces players onto a line. Baseball players rarely exist on one.

Cluster analysis lets us approach the same group differently. Instead of asking who ranks first, second, or third, we ask which players possess similar statistical profiles.

For Figure 4, we clustered the top 15 shortstops using standardized offensive performance, defensive performance, and three-season peak value.

Figure 4. Similarity Among the Top 15 Shortstops of the 1970s.

The dendrogram reinforces what Figure 2 suggested. Toby Harrah branches away from the more balanced offensive group because of his unusual offense-first profile. Concepción, Patek, and Speier cluster much more naturally with one another, while Bowa and Belanger occupy the defense-heavy side of the landscape.

This is one reason I like clustering for historical baseball analysis. Two players can finish near each other in a ranking without being remotely similar players.

The ranking tells us how much value the model identifies. The dendrogram begins to tell us what kind of player produced it.

What About Contemporary Reputation?

Only after completing the statistical rankings did I bring All-Star selections and Gold Gloves into the study.

The comparison is fascinating. Concepción made six All-Star teams during the decade and won five Gold Gloves. Larry Bowa made five All-Star teams and won two Gold Gloves. Campaneris was selected five times. Belanger, despite only one All-Star selection during the decade, accumulated seven Gold Gloves.

Concepción’s contemporary reputation therefore lines up remarkably well with the model. The statistical system identifies him as the decade’s best all-around shortstop, while the people watching and voting at the time repeatedly recognized both his overall quality and his defense.

Patek provides the more provocative contrast. He made three All-Star teams but won no Gold Gloves, yet the model places him second overall. Speier also made three All-Star teams and won no Gold Gloves during the decade, but finishes third.

Perhaps those players were better than their historical reputations now suggest. Perhaps our model is capturing forms of value that were less likely to produce awards. More likely, it is some combination of both.

Either way, the discrepancy is worth noticing.

Testing the Ranking

Any historical ranking should survive reasonable changes to its assumptions.

The original model required at least three qualifying shortstop seasons. I therefore repeated the ranking with a substantially stricter requirement of five qualifying seasons.

The top five remained unchanged: Dave Concepción, Freddie Patek, Chris Speier, Larry Bowa, and Bert Campaneris.

That stability matters. A ranking that changes dramatically because a cutoff moves from three seasons to five would make me nervous. Here, the central conclusion survives.

Concepción does not win because of an arbitrary eligibility rule. Patek and Speier do not appear near the top because of a tiny sample. Bowa and Campaneris remain firmly in the same group.

The deeper we look, the more stable the top of the ranking becomes.

What the Model Cannot Tell Us

There is still a limit to what we can claim.

Traditional fielding statistics are not modern defensive metrics. Assists depend partly on opportunity. Double plays depend partly on the second baseman and the pitching staff. Fielding percentage cannot tell us whether a defender reached a ball another shortstop would never have touched.

Those limitations become particularly important when interpreting players such as Belanger, Bowa, and Bill Russell. A ranking constructed from traditional fielding data should not be read as a precise estimate of defensive runs saved.

It is better understood as a historical reconstruction from the evidence available.

That is still valuable. The goal is not to pretend we possess Statcast data from 1975. The goal is to combine the surviving information carefully enough that patterns begin to emerge.

And several patterns do.

So, Who Was the Best?

After offense, defense, longevity, peak value, sensitivity testing, and comparison with contemporary recognition, I think the answer is Dave Concepción.

His case is unusually complete. He ranks second offensively, fifth defensively, first overall, and first in three-season peak. He also received substantial contemporary recognition, with six All-Star appearances and five Gold Gloves during the decade.

But the study would be much less interesting if it ended there.

Freddie Patek’s second-place finish may be the most surprising result. Chris Speier emerges as perhaps the most underrated balanced shortstop of the decade. Larry Bowa and Mark Belanger demonstrate the extraordinary value that defense could carry at the position. At the same time, Toby Harrah shows how far outstanding offense could push a shortstop even when his traditional defensive record lagged behind.

That is what I like most about the result.

There was no single template for a great shortstop in the 1970s. Harrah, Belanger, Bowa, Patek, Speier, Campaneris, and Concepción reached value by different paths. The data does not erase those differences by forcing everyone into the same mold.

Instead, it makes the differences visible. And when offense and defense are finally brought together, Dave Concepción is the player standing highest in the middle.

MLB Team Defense Update (9/7/26)

Defensive statistics have always been difficult to interpret.

Batting statistics usually tell a fairly direct story. A hitter gets on base, hits for power, strikes out, walks, and produces runs. Pitching statistics are more complicated, but the broad questions are still familiar. Does the pitcher miss bats? Does he limit walks? Does he prevent runs?

Defense is different.

No single defensive statistic is universally accepted as definitive. Defensive Runs Saved, Outs Above Average, Fielding Run Value, FanGraphs Def, fielding percentage, and catcher throwing statistics all attempt to measure defense from slightly different directions. Sometimes they agree. Sometimes they disagree dramatically.

That makes team defense an ideal candidate for principal component analysis. Rather than deciding in advance which defensive statistic is “correct,” PCA lets the statistics reveal the dominant patterns in the data.

For this study, I used 2026 FanGraphs team defensive data through September 7. Six measures were included:

  • Defensive Runs Saved
  • Outs Above Average
  • Fielding Run Value
  • FanGraphs Def
  • Fielding percentage
  • Caught-stealing rate

I standardized each variable before performing the PCA. This is important because the statistics exist on very different numerical scales. Fielding percentage, for example, clusters around .980 to .990, while DRS can range from strongly negative to well over +100. Without standardization, a statistic’s scale could influence the PCA more than the information it contains.

The result was surprisingly clean. The first principal component explains 60.5 percent of all variation among MLB team defenses. Even more importantly, its meaning is easy to interpret.

The loadings on PC1 were:

\mathrm{Def} = 0.507 \mathrm{FRV} = 0.506 \mathrm{OAA} = 0.493 \mathrm{DRS} = 0.393 \mathrm{FP} = 0.297

Caught-stealing rate contributed almost nothing to the first component.

In practical terms, PC1 acts as an overall defensive quality axis. That interpretation is reinforced by the extremely strong relationship between PC1 and FanGraphs Def. The correlation is about 0.97. So although PCA was not told what constituted “good defense,” it independently produced a first component that behaves almost exactly like a composite defensive-quality measure.

And one team separates itself immediately.

Chicago is not merely first; the Cubs are in a different neighborhood. Their PC1 score is 5.94, compared with 2.81 for second-place Arizona. No other team approaches Chicago’s position on the primary defensive axis. That distance is important. Rankings can sometimes exaggerate small differences. A team ranked first may be only marginally better than the team ranked second.

That is not what is happening here. The PCA suggests an enormous separation between the Cubs and everyone else.

Chicago entered September 8 with 107 Defensive Runs Saved, 65 Outs Above Average, and 64 Fielding Run Value in the data used here. Those are not merely good numbers. They represent broad agreement among different defensive measurement systems that Chicago has been exceptional.

The top ten teams by PC1 were:

Rank Team PC1
1 Chicago Cubs 5.94
2 Arizona 2.81
3 St. Louis 2.01
4 Toronto 1.81
5 San Diego 1.68
6 Los Angeles Dodgers 1.62
7 Atlanta 1.61
8 Boston 1.33
9 Kansas City 1.18
10 Cleveland 0.84

Arizona emerges as the clear second-place team. The Diamondbacks’ defensive profile is particularly strong in the advanced range-based measures. They recorded 39 OAA and 31 FRV, producing a PC1 score substantially above most of the league.

St. Louis ranks third, followed by Toronto and San Diego. Cleveland comes in tenth. That is a respectable position, but the graph makes clear how far Chicago is from a good defensive team.

The opposite end of the PCA is equally interesting. Seattle ranks last with a PC1 score of -3.29. The Athletics are close behind at -3.02, followed by the Angels, Colorado, and Minnesota.

Rank Team PC1
30 Seattle -3.29
29 Athletics -3.02
28 Los Angeles Angels -2.33
27 Colorado -2.13
26 Minnesota -1.93
25 Pittsburgh -1.82
24 Cincinnati -1.46
23 San Francisco -1.41

Seattle’s placement is driven heavily by -51 OAA and -44 FRV. Those numbers suggest a club that has struggled considerably to convert balls in play into outs relative to what would be expected.

But the lower portion of the rankings also reveals why using multiple defensive measurements is useful. Cincinnati, for example, had +1 OAA but -45 DRS. San Francisco showed almost the opposite disagreement, with +22 DRS but -17 OAA.

Which statistic should we trust? That is precisely the wrong question. Different defensive systems use different models, assumptions, opportunities, positioning adjustments, and definitions of responsibility. Disagreement among them is therefore not necessarily evidence that one system has failed.

It can also reveal uncertainty.

PCA helps by asking a different question: across all of these measurements, what common defensive signal appears most consistently? For most teams, that signal is PC1.

The second principal component tells a completely different story. PC2 explains another 17.3 percent of the total variance, bringing the first two principal components to approximately 77.9 percent of all defensive variation in the six original variables.

But PC2 is not another general-defense measure. It reflects the running game almost entirely.

The loading for caught-stealing rate on PC2 is approximately: 0.961. That is extraordinarily large.

The other variables contribute comparatively little. This means that the vertical axis in Figure 1 can essentially be interpreted as running-game control, while the horizontal axis measures broader defensive quality.

That makes several teams particularly interesting. San Diego ranks fifth overall defensively, but the Padres also sit very high on PC2. Kansas City shows a similar pattern. Their location on the graph suggests a defensive identity that differs from teams such as Arizona or Atlanta.

Those clubs may all be good defensively, but not in exactly the same way. And that is one of PCA’s greatest strengths. A simple ranking compresses every team into one number. The PCA retains structure.

Teams can be similar in overall quality while achieving that quality through different defensive profiles. However, there is an important methodological limitation.

DRS, OAA, FRV, and FanGraphs Def are not completely independent measurements. Several are derived from overlapping types of defensive information. OAA and FRV in particular are closely related, while FanGraphs Def incorporates modern defensive valuation into a broader positional framework.

Therefore, PC1 should not be interpreted as a completely new and independent defensive statistic. A better description would be a consensus advanced-defense index. That is still useful.

In fact, for this particular question, the overlap may be an advantage. If several different defensive systems all point in the same direction, PCA extracts that shared signal and gives it substantial weight.

Chicago is the clearest example. The Cubs do not rank first because one unusual statistic loves their defense. They rank first because virtually every major defensive measure agrees that their defense has been outstanding.

Seattle provides the mirror image. Several independent measurements likewise support the Mariners’ placement near the extreme negative end of PC1.

The most interesting cases may actually be the teams in between. Cincinnati and San Francisco show large disagreements among defensive systems. Pittsburgh has a positive DRS despite poor scores elsewhere. Philadelphia has relatively poor overall PC1 positioning while displaying a much stronger running-game score on PC2.

Those teams deserve additional investigation.

A natural next step also emerges.

Instead of using aggregate measures such as DRS, OAA, and Def, we can construct another PCA using the individual Fielding Run Value components: Throwing, Blocking, Framing, Arm, Range, Infield Double Plays, and First-Base Receiving.

That analysis would answer a different question. This PCA tells us who has been good and who has been bad. A component-level PCA could begin telling us why. For now, though, the broad picture is unusually clear.

Chicago has been the best defensive team in baseball through September 7, and not by a small margin. Arizona forms something of a second tier, followed by a cluster containing St. Louis, Toronto, San Diego, the Dodgers, Atlanta, and Boston.

At the other end, Seattle and the Athletics occupy the weakest part of the defensive landscape.

And perhaps most importantly, the analysis demonstrates why PCA is so useful in baseball analytics.

Defense does not have to be reduced to a debate over which statistic is best. Sometimes the better approach is to let the statistics vote.

 

Which MLB Teams Are Actually Pitching the Best in 2026? (A Team Pitching Study Through September 4, 2026)

ERA is useful. It is also incomplete.

A team can post an excellent ERA because its pitchers dominate hitters, limit hard contact, avoid walks, and miss bats. But a good ERA can also reflect defense, sequencing, favorable outcomes with runners on base, or simply a stretch in which balls have found gloves instead of grass.

The reverse can happen too. A pitching staff can do many of the things we associate with good pitching and still carry an ERA that makes it look merely average.

That distinction becomes particularly interesting when we look across all 30 major-league teams in 2026.

Using FanGraphs team pitching data through September 4, I wanted to answer a slightly different question from the usual one. Instead of asking which teams have allowed the fewest earned runs, I asked:

Which teams appear to be pitching the best underneath those results?

That question leads us to Philadelphia.

But it also leads us to Milwaukee, Cleveland, Atlanta, Arizona, the Yankees, and several teams whose records do not line up neatly with the quality of their pitching.

Measuring the Pitching Process

No single pitching statistic completely describes a staff.

ERA tells us what happened. FIP focuses heavily on strikeouts, walks, hit batters, and home runs. xFIP normalizes the home-run component. SIERA attempts to estimate run prevention while accounting for strikeouts, walks, and batted-ball tendencies. xERA uses Statcast information about the quality of contact allowed.

Then there is K-BB%, one of the cleanest measures of pitcher control.

A staff that strikes out many hitters while walking few is usually doing something right.

I therefore built an Underlying Pitching Process Index using seven measures:

\mathcal{M} = \left\{ \mathrm{xERA}, \mathrm{FIP}^{-}, \mathrm{xFIP}^{-}, \mathrm{SIERA}, \mathrm{K\!-\!BB\%}, \mathrm{Barrel\%}, \mathrm{HardHit\%} \right\}

The first step was to standardize every statistic across the 30 teams.

For team and pitching metric :

z_{i,m} = \frac{ x_{i,m} - \overline{x}_m }{ s_m }

There is an important complication.

Higher K-BB% is good. Higher xERA is not. The same problem applies to FIP-, xFIP-, SIERA, Barrel%, and HardHit%.

I therefore adjusted the direction of each standardized statistic so that higher always means better pitching:

q_{i,m} = d_m z_{i,m}

where:

d_m = \begin{cases} +1, & \mathrm{if\ higher\ values\ are\ better} \\ -1, & \mathrm{if\ lower\ values\ are\ better} \end{cases} q_{i,m} = d_m z_{i,m}

where:

d_m = \begin{cases} +1, & \text{if higher values are better} \\ -1, & \text{if lower values are better} \end{cases}

The seven adjusted z-scores were then averaged:

P_i = \frac{1}{7} \sum_{m \in \mathcal{M}} q_{i,m}

Finally, I converted that score into an index centered on 100, with a standard deviation of 15:

I_i = 100 + 15 \left( \frac{ P_i - \overline{P} }{ s_P } \right)

A score of 100 represents approximately league-average underlying pitching. A score of 115 is about one standard deviation above average.

This is important: the Process Index is not a FanGraphs statistic. It is a composite measure I constructed for this study from FanGraphs data.

And one team separates itself immediately.

Philadelphia Comes Out on Top

Figure 1 ranks all 30 teams using the Underlying Pitching Process Index.

Figure 1. 2026 MLB Team Pitching Through September 4: Underlying Pitching Process Ranking.

The index combines xERA, FIP-, xFIP-, SIERA, K-BB%, Barrel% allowed, and HardHit% allowed. The MLB average is centered at 100.

Philadelphia ranks first with a Process Index of 131.2.

Milwaukee is close behind at 128.6. The Yankees rank third, followed by Boston and the Dodgers.

The top ten are:

Rank Team Process Index
1 Philadelphia 131.2
2 Milwaukee 128.6
3 New York Yankees 121.6
4 Boston 118.6
5 Los Angeles Dodgers 117.4
6 Toronto 110.2
7 Cleveland 109.3
8 New York Mets 108.0
9 Pittsburgh 107.3
10 Detroit 107.3

Philadelphia’s position may initially seem strange.

The Phillies have a 4.00 ERA. Nobody looking only at ERA would identify that as the performance of baseball’s best pitching staff.

The underlying numbers tell a different story.

Philadelphia has a 3.56 xERA, 3.76 FIP, 3.54 xFIP, and 3.50 SIERA. Most impressively, the Phillies have a 17.9% K-BB%, the best figure in baseball through September 4.

That combination is difficult to dismiss.

They miss bats. They limit walks. Their contact profile is excellent. Several statistics designed specifically to isolate pitcher performance show a staff substantially better than its 4.00 ERA.

Milwaukee presents a different case.

The Brewers rank second in underlying process, with an index of 128.6, but their actual ERA is already excellent at 3.50. Their xERA is 3.47.

For Milwaukee, there is very little disagreement between process and outcome.

That distinction becomes important.

ERA and xERA Do Not Always Tell the Same Story

One simple way to measure the difference between observed and expected run prevention is:

\Delta \mathrm{ERA}_i = \mathrm{ERA}_i - \mathrm{xERA}_i

A positive value means the team’s ERA has been worse than its xERA. A negative value means the team has allowed fewer earned runs than its expected ERA would suggest.

Figure 2 shows the relationship directly.

Figure 2. 2026 Team ERA vs. xERA Through September 4.

Teams below the diagonal have an ERA lower than their xERA. Teams above the diagonal have an ERA higher than expected from their Statcast contact profile.

Philadelphia stands out.

For the Phillies:

\Delta \mathrm{ERA}_{\mathrm{PHI}} = 4.00 - 3.56 = 0.44

Their ERA is approximately 0.44 runs higher than their xERA.

That is a meaningful gap over nearly an entire season.

The Athletics are the extreme example of the opposite kind of outcome. Their ERA is 5.47, while their xERA is 4.48:

\Delta \mathrm{ERA}_{\mathrm{ATH}} = 5.47 - 4.48 = 0.99

Oakland’s underlying pitching is not good. The Athletics rank only 24th in the Process Index.

But a 5.47 ERA makes the staff appear even worse than its underlying performance suggests.

Now look below the diagonal.

Arizona has a 4.17 ERA against a 4.72 xERA:

\Delta \mathrm{ERA}_{\mathrm{ARI}} = 4.17 - 4.72 = -0.55

Atlanta is similar:

\Delta \mathrm{ERA}_{\mathrm{ATL}} = 3.57 - 4.04 = -0.47

The Cubs sit at 4.16 versus a 4.62 xERA, while the Yankees have produced an exceptional 3.23 ERA despite a 3.67 xERA.

Those differences do not mean that one statistic is right and the other is wrong.

They tell us that something interesting is happening between underlying pitcher performance and actual runs allowed.

Process Is Not the Same as Results

To examine that distinction more directly, I created a second index.

The Results Index uses ERA-, RA9-WAR, and WPA. Unlike the Process Index, these measures deliberately capture more of what actually happened on the field.

The underlying results score is:

R_i = \frac{1}{3} \left( -z_{i,\mathrm{ERA}^{-}} + z_{i,\mathrm{RA9\!-\!WAR}} + z_{i,\mathrm{WPA}} \right)

It is then converted to the same 100-centered scale:

J_i = 100 + 15 \left( \frac{ R_i - \overline{R} }{ s_R } \right)

Figure 3 compares the two indices.

Figure 3. 2026 Team Pitching: Underlying Process vs. Actual Results.

Teams above the diagonal have obtained better results than their underlying pitching process would predict. Teams below the diagonal have performed worse.

The diagonal is the key.

Teams close to it have received results consistent with the quality of their underlying pitching. Teams far above or below it deserve more attention.

Philadelphia ranks first in process but only eighth in results.

Its Results Index is almost 20 points below its Process Index.

Arizona goes the other way. The Diamondbacks rank only 25th in underlying process, but 11th in results.

Atlanta is another striking example. The Braves rank 12th in process but second in results.

The Cubs rank 26th in process and 17th in results.

These teams have converted their pitching performances into actual run prevention more effectively than the underlying indicators alone would predict.

Philadelphia has not.

Neither have the Athletics, the Seattle Mariners, the San Francisco Giants, or the Mets.

There are many possible reasons. Defense is a factor. So does sequencing. Relievers inherit runners. Pitchers behave differently with runners on base. Ballparks are important as is random variation.

The point is not that all deviation must eventually disappear.

It is that ERA alone completely hides the deviation.

Strikeouts, Walks, and SIERA

One of the strongest relationships in the data appears when we compare K-BB% with SIERA.

Figure 4. 2026 Team Command and Dominance Through September 4.

Teams toward the right strike out more hitters relative to walks. Teams toward the top have lower SIERA.

K-BB% is appealing because it strips pitching down to two outcomes over which the pitcher exercises considerable control.

The calculation is simple:

\mathrm{K\!-\!BB\%} = \mathrm{K\%} - \mathrm{BB\%}

Philadelphia again occupies elite territory.

The Phillies lead MLB at 17.9%. Milwaukee follows at 17.3%. The Dodgers are at 16.5%, with the Yankees and Cleveland both around 16%.

This is one reason Cleveland deserves more attention.

The Guardians are not surviving through a suspiciously low ERA or an unusually favorable gap between ERA and expected statistics. They rank fifth in SIERA and fifth in K-BB%.

That is a much stronger foundation.

Cleveland’s Pitching Is Not the Problem

Cleveland entered September 5 with a 72-70 record.

The Guardians’ pitching numbers, however, look like those of a considerably better team.

Cleveland ranks seventh in ERA, eighth in xERA, sixth in FIP, sixth in xFIP, fifth in SIERA, fifth in K-BB%, seventh in pitching WAR, seventh in the Underlying Pitching Process Index, and seventh in the Results Index.

That consistency is important.

Cleveland is not seventh because one unusual statistic dragged the composite upward. Almost every important run-prevention and defense-independent measure puts the Guardians somewhere in the same general neighborhood.

Their ERA is 3.75.

Their xERA is 3.90.

Their FIP is 3.86, xFIP is 3.88, and SIERA is 3.75.

There is, however, one weakness.

Cleveland ranks only around 21st in Barrel% allowed and 22nd in HardHit% allowed. Opponents have made better contact against the Guardians than the overall pitching rankings might imply.

They have compensated for that weakness largely through strikeouts and control.

That creates a somewhat unusual pitching profile. Cleveland is not especially dominant at suppressing every kind of hard contact. Still, the Guardians prevent enough balls from being put into play in the first place and avoid giving away enough free bases to remain an excellent overall staff.

If the question is why Cleveland has been only a little above .500, the team’s pitching data strongly suggests looking elsewhere.

How Much Does Pitching Explain Winning?

Of course, a baseball team does not win with pitching alone.

To measure the association between underlying pitching quality and the standings, I compared the Process Index with team winning percentage.

Winning percentage is:

\mathrm{WinPct}_i = \frac{ W_i }{ W_i + L_i }

I then estimated a simple linear regression:

\mathrm{WinPct}_i = \alpha + \beta I_i + \varepsilon_i

The resulting relationship appears in Figure 5.

Figure 5. 2026 Underlying Pitching Quality vs. Team Winning Percentage.

The regression compares each team’s Process Index with its winning percentage through September 4.

The coefficient of determination is:

R^2 = 0.314

Approximately 31.4% of the cross-team variation in winning percentage is associated with variation in the Underlying Pitching Process Index in this 2026 snapshot.

Pitching clearly is essential, but nearly 69% of the variation remains elsewhere.

The Cubs are 80-62 despite ranking 26th in underlying pitching process. Atlanta is 84-57 while ranking only 12th. Cleveland is 72-70 despite ranking seventh.

A team can compensate for mediocre pitching.

It can also waste very good pitching.

Philadelphia May Be Better Than Its ERA

The Phillies may be the most important example in the entire study.

Their 4.00 ERA does not look dominant. If we stopped there, we would never consider Philadelphia the best pitching team in baseball.

Yet the deeper indicators repeatedly return to the same conclusion.

Philadelphia ranks first in the Process Index. The Phillies lead baseball in K-BB%. Their xERA is 3.56. Their xFIP is 3.54. Their SIERA is 3.50.

Their actual ERA is the outlier.

That does not guarantee that the ERA will suddenly collapse toward 3.50. Baseball statistics do not work like a mechanical spring that must snap back to equilibrium.

But if I were trying to determine which pitching staff I trusted going forward, I would place more weight on the cluster of underlying measures than on ERA alone.

Philadelphia’s process has been elite.

Milwaukee Has the Cleaner Case

Milwaukee requires less explanation.

The Brewers rank second in the Process Index at 128.6 and have a 3.50 ERA against a 3.47 xERA.

There is almost no gap.

Their process says they are excellent. Their results say they are excellent. Their 88-54 record agrees.

Sometimes the complicated analysis confirms the obvious.

That is useful too.

Milwaukee does not need a regression argument or a discussion of sequencing to explain its success. The Brewers have simply pitched extremely well.

The Yankees Have Turned Good Pitching Into Great Results

The Yankees provide another variation.

New York ranks third in underlying process but first in the Results Index.

Their 3.23 ERA is substantially better than their 3.67 xERA:

\Delta \mathrm{ERA}_{\mathrm{NYY}} = 3.23 - 3.67 = -0.44

The underlying staff is genuinely good. This is not a weak pitching team disguising itself behind a low ERA.

But the results have been even better than the already strong process.

That distinction is noteworthy.

The Yankees do not appear to be a mirage. They appear to be an excellent pitching staff that has also benefited from a gap between underlying performance and runs actually allowed.

Arizona Is More Difficult to Trust

Arizona presents a much different problem.

The Diamondbacks’ 4.17 ERA is not particularly impressive on its own, but it looks much better compared to their 4.72 xERA.

Their Process Index ranks only 25th in baseball.

Their Results Index ranks 11th.

That is one of the largest process-to-results gaps in the league.

Perhaps Arizona can continue converting that underlying performance into acceptable run prevention. There may be legitimate reasons for part of the difference.

Still, if the goal is prediction rather than description, I would be considerably more cautious about Arizona than Philadelphia.

The Phillies have poor results relative to a strong process.

Arizona has strong results relative to a poor process.

Those are not equivalent situations.

Oakland Shows How Ugly Results Can Become

At the bottom end, Oakland offers an equally useful lesson.

The Athletics rank 24th in underlying pitching process. That is bad, but not the worst in baseball.

Their Results Index ranks 30th.

Their 5.47 ERA is almost one full run higher than their 4.48 xERA.

The underlying numbers therefore suggest two conclusions at once.

Oakland has not pitched well.

Oakland has also gotten even worse results than its mediocre pitching would normally lead us to expect.

Both statements can be true.

What the Rankings Really Tell Us

This study is not intended to replace ERA with another single magic number.

In fact, that would defeat the purpose.

The more interesting lesson is that pitching has several layers.

There is the process: missing bats, limiting walks, controlling contact, suppressing barrels, and producing outcomes that FIP, xFIP, SIERA, and xERA consider sustainable.

Then there are the results.

Usually they move together.

Sometimes they do not.

Philadelphia has arguably produced the strongest underlying pitching process in baseball, yet received only the eighth-best composite results.

Milwaukee has been excellent by both standards.

The Yankees have turned excellent underlying pitching into even better results. Atlanta and Arizona have substantially outperformed what their process metrics would lead us to expect.

And Cleveland?

Cleveland may be the most revealing team of all.

The Guardians are barely above .500, yet almost every pitching measure says the staff belongs among the top quarter of baseball. Their Process Index ranks seventh. Their Results Index also ranks seventh.

There is little evidence that Cleveland’s pitching has betrayed the team.

If anything, it has kept the Guardians afloat.

Final Thoughts

ERA remains one of baseball’s most intuitive statistics because it measures something real. Runs crossed the plate, and those runs counted.

But describing what happened is not always the same as understanding why it happened.

That is where the deeper statistics become valuable.

Through September 4, Philadelphia appears to have baseball’s strongest underlying team pitching process. Milwaukee is close behind and has translated that process into superior run prevention. The Yankees have done even better in converting strong pitching into actual results.

Cleveland quietly belongs in the next group.

Meanwhile, Atlanta, Arizona, and the Cubs remind us that actual run prevention can exceed expectations based on underlying statistics, sometimes by a considerable margin.

None of those observations requires us to choose between ERA and advanced metrics.

We can use both.

ERA tells us what the scoreboard recorded.

The underlying statistics help us understand how the pitchers got there and perhaps where they are going next.

That difference is where the interesting part begins.

 

Who Is Producing the Most Offense in Baseball in 2026 Through Early September?

We tend to talk about hitting leaderboards as though they answer a simple question. Who has been the best hitter?

It is not quite that simple. A hitter can produce enormous value through power, another through reaching base, another through a combination of contact and baserunning. Rate statistics can identify extraordinary performance while overlooking playing time. Counting statistics reward accumulation but can obscure efficiency. WAR goes even further, mixing offense with defense and positional value.

For this study, I wanted to ask a narrower question: Who has actually produced the most offensive value in Major League Baseball so far in 2026?

I used my September 5 FanGraphs downloads, containing 142 hitters and a wide range of traditional, advanced, Statcast, baserunning, and value statistics. The results reveal a clear No. 1. They also reveal something more interesting underneath.

Measuring offensive production

The primary statistic for the study is FanGraphs Offense. In the data, it has a particularly useful decomposition:

\mathrm{Offense}_i = \mathrm{Batting}_i + \mathrm{BaseRunning}_i

That makes it attractive for this question.

Defense is absent. Positional adjustment is absent. We are measuring what a player has contributed when his team is at the plate or when he is running the bases.

WAR is still useful, but it answers a different question. Here I wanted to isolate offense.

The initial leaderboard is striking.

Yordan Alvarez leads baseball with 52.31 offensive runs above average. Pete Crow-Armstrong is very close behind at 49.45.

Then comes a cliff.

James Wood is third at 33.68, followed by Shohei Ohtani at 32.85 and Randy Arozarena at 30.61. The difference between Alvarez and Crow-Armstrong is only 2.86 runs. The difference between Crow-Armstrong and Wood is nearly 15.8 runs.

So, at least by this measure, 2026 has developed a clear top tier of two.

Alvarez and Crow-Armstrong arrived there differently

The decomposition is useful because Alvarez and Crow-Armstrong have not constructed their offensive value in the same way.

For Alvarez:

\mathrm{Offense}_{\mathrm{Alvarez}} = 57.01 - 4.70 = 52.31

His bat has been so productive that he can lose almost five runs through baserunning and still lead everyone.

Crow-Armstrong presents almost the opposite profile:

\mathrm{Offense}_{\mathrm{CrowArmstrong}} = 43.77 + 5.68 = 49.45

His batting contribution is excellent, but his legs add another 5.68 runs.

That distinction can be seen clearly when batting and baserunning are separated.

Alvarez is sitting far to the right because of his extraordinary batting value, but below zero in baserunning. Crow-Armstrong combines one of baseball’s best bats with strongly positive running value.

James Wood and Ohtani fall somewhere between those extremes.

There is no single recipe.

Rate production tells almost the same story

A natural objection to using total offensive runs is that playing time matters. A player can accumulate more value simply because he has received more plate appearances.

That is where wRC+ becomes useful.

Across the 142 hitters in the dataset, FanGraphs Offense and wRC+ have a correlation of approximately:

r = 0.970

A simple regression gives:

\widehat{\mathrm{Offense}}_i = -60.98 + 0.618 \left( \mathrm{wRC}^{+}_i \right)

with:

R^2 = 0.941

That is an exceptionally tight relationship.

It should not be interpreted as independent validation, since Offense and wRC+ are built from related offensive information. What it does show is that differences in playing time and baserunning have not radically reordered the best hitters. The strongest rate producers are generally producing the most total offensive value as well.

And once again, Alvarez sits at the top.

His 180.5 wRC+ is the best in the dataset. Crow-Armstrong follows at 157.6, while Willson Contreras, James Wood, Bryce Harper, Randy Arozarena, Junior Caminero, and Ohtani form the next group.

An approximately 181 wRC+ means that Alvarez has created runs at roughly 81 percent above league average, after the adjustments built into wRC+.

That is a remarkable offensive season.

But what should have happened?

Observed production is only half of the story.

Statcast gives us another way to look at these hitters. Instead of merely asking what happened, xwOBA asks what we would expect from the quality of the contact, along with the other inputs incorporated into the statistic.

I calculated the difference as:

\Delta \mathrm{wOBA}_i = \mathrm{wOBA}_i - \mathrm{xwOBA}_i

A positive value means the player’s actual wOBA has exceeded his xwOBA. A negative value indicates that his expected production is higher than his observed production.

The comparison changes the story.

Alvarez has a .4289 wOBA, which is already the best in the sample.

His xwOBA is .4496.

\Delta \mathrm{wOBA}_{\mathrm{Alvarez}} = 0.4289 - 0.4496 = -0.0207

In other words, the underlying contact data do not suggest that Alvarez has been getting lucky.

They suggest the opposite.

His offensive production has actually fallen about 21 points of wOBA below what Statcast would expect from the underlying events.

Crow-Armstrong provides a fascinating contrast.

\Delta \mathrm{wOBA}_{\mathrm{CrowArmstrong}} = 0.4013 - 0.3684 = 0.0329

His actual wOBA has exceeded his xwOBA by roughly 33 points.

That does not invalidate what he has accomplished. Runs that have already scored still count. It does suggest, however, that Alvarez’s underlying offensive performance is considerably stronger than the small difference between their total Offense numbers might initially imply.

James Wood might be the most frightening hitter behind Alvarez

Wood ranks third in total Offense at 33.68 runs, but that ranking almost understates what is happening.

His xwOBA is .4183, second only to Alvarez.

His barrel rate is 20.6 percent, the highest in the dataset. His hard-hit rate is 58.1 percent, also the highest.

Then there is his walk rate: 16.7 percent.

That is an unusual combination. Wood is not simply crushing baseballs. He is also refusing pitches frequently enough to force pitchers into the strike zone.

His main weakness remains strikeouts. His K% is 28.6 percent, which is high. Yet when he connects, the quality of contact is extraordinary.

The upper-right portion of Figure 5 is where we want to look. That region combines high barrel frequency with high expected offensive production.

Wood is there.

So is Alvarez, although Alvarez achieves his even higher xwOBA without matching Wood’s astonishing barrel percentage. Ohtani also occupies the elite region, while Pete Alonso combines tremendous hard contact with a .3946 xwOBA.

These are different offensive machines producing similar outcomes.

Bobby Witt Jr. is an important counterexample

Bobby Witt Jr. does not appear among the very top total offensive producers. His Offense value is 20.88, which ranks 21st in the dataset, and his wRC+ is 121.5.

Yet his xwOBA is .3820.

His actual wOBA is only .3488.

\Delta \mathrm{wOBA}_{\mathrm{Witt}} = 0.3488 - 0.3820 = -0.0332

That is one of the largest negative gaps in the entire sample.

Witt has also contributed 7.36 baserunning runs, third best among these hitters. His overall WAR is 6.31, second only to Crow-Armstrong in the dataset.

This is precisely why the definition of “best” matters.

If we ask who has produced the most offense, Witt is not at the very top. If we ask who has played the most valuable all-around baseball, the answer changes substantially. If we ask whose contact suggests better offensive results than he has received, Witt suddenly becomes extremely interesting.

One leaderboard cannot answer all three questions.

A standardized look at offensive styles

To compare very different statistics on a common scale, I standardized six measures across the full 142-player sample.

For player and metric :

z_{i,m} = \frac{ x_{i,m} - \overline{x}_m }{ s_m }

For strikeout rate, I reversed the sign so that positive values always represent the favorable direction:

z^{*}_{i,\mathrm{K\%}} = - z_{i,\mathrm{K\%}}

The resulting profiles are shown below.

The contrast is useful.

Alvarez is approximately 3.5 standard deviations above the sample mean in wRC+ and almost 3.8 standard deviations above average in xwOBA. There is no obvious weakness at the plate. His major negative component comes from baserunning.

Crow-Armstrong’s profile is more balanced between hitting and running. His ISO is exceptional, and his baserunning value is roughly two standard deviations above the sample mean, but his xwOBA is much less extreme than his actual offensive results.

Wood has perhaps the most intriguing shape of all. His xwOBA, power, and walk rate are spectacular, while his strikeout rate pulls strongly in the opposite direction.

Bryce Harper shows yet another profile. He combines elite plate discipline with strong expected production, but negative baserunning reduces his total contribution.

Four hitters survive almost every test

One way to avoid becoming overly dependent on a particular metric is simply to ask which players remain near the top regardless of how we look.

Only four hitters rank in the top 10 in Offense, wRC+, and xwOBA:

Player Offense Rank wRC+ Rank xwOBA Rank
Yordan Alvarez 1 1 1
James Wood 3 4 2
Shohei Ohtani 4 8 3
Bryce Harper 8 5 5

That is a formidable group.

But even within this smaller group, Alvarez separates himself. He is not merely first by one convenient statistic. He is first in total Offense, first in wRC+, first in wOBA, and first in xwOBA.

He also leads this dataset in WPA and RE24.

Different approaches keep returning the same name.

So who has been the best offensive player?

Through September 5, I think the answer is Yordan Alvarez, and the case is unusually strong.

Pete Crow-Armstrong has been close in actual total offensive value and has added considerably more on the bases. His season is extraordinary. But once expected production is introduced, the small gap between them begins to look larger. Alvarez’s .450 xwOBA suggests that the quality of his offensive performance may be even better than the already spectacular results indicate.

James Wood deserves special attention as well. If the question changes from “Who has produced the most?” to “Whose underlying offensive profile frightens me most going forward?”, Wood becomes a serious candidate. A 20.6 percent barrel rate, 58.1 percent hard-hit rate, .418 xwOBA, and 16.7 percent walk rate is an extraordinary collection of traits.

Ohtani and Harper complete what might reasonably be called the most robust elite group in the data.

The larger lesson, however, is about measurement itself. Statistics do not merely rank players. Different statistics ask different questions.

Offense tells us how much was produced. wRC+ tells us how efficiently it was produced. xwOBA tells us what the underlying contact suggests should have been produced. Baserunning shows us how value can be created after the ball leaves the bat.

When all of those measurements point toward the same player, the conclusion becomes difficult to avoid.

So far in 2026, Yordan Alvarez has been baseball’s most impressive offensive producer. That is a fact.

 

Stem-and-Leaf Plots and Histograms: Two Views of the Same Data

I downloaded 2026 MLB OPS data through the end of August to create the following figures. I always start with Stem-and-Leaf Plots when I begin a study. This short post is not about the distribution of OPS in Major League Baseball; I simply want to illustrate the difference between Stem-and-Leaf Plots and Histograms. Why? I think it is a useful exercise, especially for aspiring scientists or anyone curious about how to approach data.

A histogram and a stem-and-leaf plot can look remarkably similar. Both are designed to reveal the distribution of a numerical variable. Both can show where observations cluster, where the tails extend, and whether the distribution appears symmetric or skewed.

But they do not show the data in quite the same way.

Consider the OPS values for the 142 players in this dataset. The values range from .542 to 1.035, with a mean of .756 and a median of .747. When those values are placed in a histogram, the overall shape of the distribution becomes immediately visible.

The histogram excels at this.

Grouping OPS values into intervals provides a quick visual summary of where most players are concentrated. We can see the center of the distribution, its spread, and the relatively small number of players occupying the extreme upper and lower ends. If the primary question is “What does this distribution look like?”, the histogram is difficult to beat.

There is a cost, however. The individual observations disappear.

Suppose a histogram bar represents OPS values between .750 and .775. We know how many players fall within that interval, but we cannot see their actual OPS values. A player with a .751 OPS and another with a .774 OPS are simply members of the same bin.

The stem-and-leaf plot preserves that information.

Using a key such as

0.7 | 5 = .75 OPS

we can still see the shape of the distribution, but the leaves retain the underlying observations. A cluster of leaves tells us that many players occupy a particular region, while the individual digits allow us to reconstruct their approximate OPS values.

The split-stem version offers an especially useful compromise. Each tenth is divided into two rows, with leaves 0 through 4 on one row and 5 through 9 on the next. It creates more visual separation than the highly condensed stem-and-leaf plot, without producing the very long display that results from using hundredths as individual stems.

This highlights an important distinction between the two techniques.

A histogram emphasizes shape.

A stem-and-leaf plot emphasizes shape while preserving the data.

That makes the histogram particularly effective for larger datasets and quick visual comparisons. The stem-and-leaf plot is often more revealing with small or moderate datasets because it allows us to move back and forth between the distribution and the observations that created it.

Neither graph is inherently better.

They answer slightly different questions.

The histogram asks us to look at the forest. The stem-and-leaf plot lets us see the forest while still being able to identify many of the trees.

For exploratory data analysis, there is considerable value in looking at both.

POSTSCRIPT

I took a closer look at the two plots from above. The most interesting thing about them is the extreme outlier evident at the higher end. This is exactly what Exploratory Data Analysis is for. The outlier is Yordan Alvarez of the Houston Astros. Take a look at the Box Plot I created based on the same data as the Stem-and-Leaf Plot and Histogram.

The small circle on the right-hand side represents Alvarez. He is having an exceptional season; his OPS is a bit of an anomaly. His performance is far above what everyone else in the league has achieved. In this instance, the Box Plot is my favorite visualization.  It clearly shows how unusual Alvarez has been this year.

 

The All-Star Who (Initially) Did Not Look Like One

The All-Star Who (Initially) Did Not Look Like One

Did Travis Bazzana Really Deserve His 2026 Selection?

When Travis Bazzana was named to the 2026 American League All-Star team, my first reaction was surprise. I mean, I was shocked. I didn’t expect him to make the team, did you?

Bazzana had certainly been good. But an All-Star already? He had not even reached the major leagues until April 28. By the time the All-Star roster was selected, he had played only 58 games and accumulated 249 plate appearances.

That made the selection worth examining. I really want to know how and why he made the team.

There was another reason to look closely. Bazzana was not simply added by Major League Baseball to make sure Cleveland had a representative. MLB’s official roster announcement specifically identified him as the American League’s second baseman selected through the Player Ballot, a vote involving players, managers, and coaches. Ernie Clement of Toronto had already been elected the starting second baseman by the fans.

So this was not merely a question of whether Bazzana was good. The more interesting question was whether the players themselves got it right.

Defining “Deserved”

There are several subtly different ways to ask whether someone deserved to be an All-Star. The first is simple: Was Bazzana playing at an All-Star level?

The second is more demanding: Was Bazzana one of the best American League second basemen at the time the roster was selected?

And the third is the hardest: Was he the most deserving reserve once Clement had already been elected the starter?

Those questions do not necessarily produce the same answer. A player can be performing at an All-Star level without being the single best choice available. I’ll take a look to see if he was the best choice.

For this analysis, I used the FanGraphs data available through July 4, immediately before the roster announcement. Clearly, using August statistics to judge a July decision would allow information that voters could not have possessed at the time to contaminate the analysis.

I also imposed a 150-plate-appearance minimum. That eliminated tiny samples while comfortably retaining Bazzana’s 249 plate appearances.

One player required special treatment. Ezequiel Duran did not appear in the FanGraphs second-base-filter export, but he was certainly an official AL second-base candidate on MLB’s ballot and appeared in the broader qualified American League FanGraphs exports. MLB’s June ballot updates even showed him among the leaders at the position. I therefore restored him to the comparison pool.

The First Surprise: Bazzana Was Good

The raw FanGraphs numbers immediately undermine the argument that Bazzana was a reputation-driven selection.

He was slashing:

.250/.341/.412

with 7 home runs, 12 stolen bases, an 11.6 percent walk rate, and a 20.9 percent strikeout rate.

More importantly, his 114 wRC+ indicates that his offensive production was roughly 14 percent better than league average after the adjustments incorporated into the metric. He had accumulated +6.15 offensive runs, +2.07 baserunning runs, and 1.43 WAR in only 249 plate appearances. Those are legitimate numbers.

They do not immediately prove that Bazzana should have been the reserve second baseman. But they establish something important at the outset. The selection was not absurd.

Table 1. Leading AL Second-Base Candidates Through July 4, 2026

Player PA wRC+ BsR Def WAR
Ezequiel Duran 297 107 +0.88 +8.00 2.22
Jazz Chisholm Jr. 334 98 +4.67 +4.83 2.10
Chase Meidroth 364 104 -0.29 +4.36 1.90
Kody Clemens 301 123 +0.68 -2.22 1.76
Travis Bazzana 249 114 +2.07 -0.91 1.43
Ernie Clement 342 108 -0.60 -1.08 1.38
Cole Young 357 107 -0.81 -1.37 1.37
Gleyber Torres 190 128 -2.91 -0.18 1.01

The FanGraphs exports place Bazzana behind Duran, Jazz Chisholm Jr., Chase Meidroth, and Kody Clemens in total WAR, but slightly ahead of Clement and Cole Young. Duran’s broader AL export shows 2.22 WAR, driven in considerable part by a very strong defensive rating.

Figure 1 changes the tone of the discussion. Figure 1 shows the FanGraphs WAR totals for the leading American League second-base candidates through July 4. The important point is not simply that Bazzana ranked fifth. He remained within the main cluster of serious candidates despite having considerably fewer plate appearances than most of the players ahead of him.

Bazzana was not the WAR leader. He was not even particularly close to Duran. Yet neither was he buried among mediocre players. He occupied the middle of a fairly compact cluster of serious candidates.

That is our first important result.

The Offensive Case for Bazzana

If we look only at offense, Bazzana’s case becomes considerably stronger.

Kody Clemens had the best combination of power and overall offensive production among the leading candidates, carrying a 123 wRC+ and .500 slugging percentage. Gleyber Torres produced an even higher 128 wRC+, although he had only 190 plate appearances and lost substantial value on the bases.

Bazzana sat immediately behind that offensive tier.

His .341 on-base percentage was particularly valuable. He was walking frequently, making enough contact, adding modest power, and creating additional value with his legs.

FanGraphs decomposes offensive value in a useful way:

\mathrm{Offense} = \mathrm{Batting} + \mathrm{Base\ Running}

For Bazzana:

\mathrm{Offense}_{\mathrm{Bazzana}} = 4.076 + 2.074 = 6.150

That is a strong total for someone with only 249 plate appearances. FanGraphs credited him with roughly 4.1 batting runs and another 2.1 runs on the bases.

The baserunning component should not be dismissed as decorative. Twelve stolen bases in 58 games certainly attracted attention, but BsR attempts to measure more than stolen bases alone. Bazzana was producing real value outside the batter’s box.

This is where his All-Star case begins to make intuitive sense. Players facing Cleveland were not necessarily thinking about Bazzana’s WAR ranking. They were seeing a hitter who controlled the strike zone, got on base, ran well, and was already producing above-average offense only weeks into his major-league career.

The Player Ballot may have been capturing something real.

Then Defense Changes Everything

The reason Bazzana’s total WAR was not higher was defense. FanGraphs credited him with -0.91 Def through the cutoff date. That figure combines fielding performance with positional value.

Conceptually:

\mathrm{Defense} = \mathrm{Fielding} + \mathrm{Positional}

For Bazzana, the underlying FanGraphs components were approximately -1.27 fielding runs and +0.36 positional runs, producing the -0.91 total. Compare that with Duran. Duran had only a 107 wRC+, lower than Bazzana’s 114. His total offensive value was +3.31 runs, also well below Bazzana’s +6.15.

But Duran’s FanGraphs defense was an extraordinary +8.00 runs. That defensive advantage propelled him to 2.22 WAR.

Jazz Chisholm Jr. followed a similar path. His 98 wRC+ was actually below league average, yet his baserunning and defense were excellent. FanGraphs gave him +4.67 BsR and +4.83 Def, enough to reach 2.10 WAR.

Kody Clemens was almost the mirror image. His offense was excellent, but FanGraphs rated his defense at -2.22 runs.

That produces an unusually interesting positional race.

Figure 2 plots offensive production against FanGraphs defensive value and shows just how different the candidates were. Bazzana sits on the offense-heavy side of the group, while Duran, Chisholm, and Meidroth received substantially more help from defense.

There was no single prototype for a valuable American League second baseman in early 2026.

Duran and Chisholm were being pushed upward by defense. Kody Clemens was being driven by his bat. Bazzana was getting most of his value from batting and baserunning while surrendering a relatively small amount defensively.

Figure 3 makes that contrast even clearer. The same contrast becomes more evident when the components are separated. Figure 3 compares batting, baserunning, and defensive value for the principal candidates. Bazzana’s profile is distinctive: strong batting value, meaningful baserunning value, and a relatively small defensive penalty.

This is also why defense must be included in the analysis. Ignoring it would artificially elevate Bazzana and Kody Clemens while punishing Duran, Chisholm, and Meidroth.

At the same time, there is a legitimate reason not to treat 80 games of defensive measurement as absolute truth. Defensive statistics stabilize slowly. Positioning, opportunity, scoring systems, and relatively small numbers of fielding plays can substantially affect estimates.

That creates a useful sensitivity test.

What Happens If We Trust Defense Less?

Instead of pretending that defensive measurement is perfectly precise, we can ask how the ranking changes as we gradually reduce its influence.

FanGraphs’ Runs Above Replacement framework can be represented approximately as:

\mathrm{RAR} = \mathrm{Offense} + \mathrm{Defense} + \mathrm{League} + \mathrm{Replacement}

WAR then converts those runs into wins:

\mathrm{WAR} = \frac{ \mathrm{RAR} }{ \mathrm{Runs\ per\ Win} }

For a sensitivity check, I recalculated the candidates while allowing defense to receive 0, 0.5, or its full FanGraphs weight. This is not a replacement version of WAR. It is simply a way to see how dependent the ranking is on defensive evaluation. Because defensive measurements are especially uncertain over partial seasons, it is worth asking whether the conclusion depends heavily on them. Figure 4 performs that sensitivity test by recalculating the candidates’ diagnostic value with defense given zero weight, half weight, and its full FanGraphs weight.

The result is revealing. When defense receives its full FanGraphs weight, Duran leads at 2.22 WAR, followed by Chisholm and Meidroth. Bazzana sits at 1.43.

When defense is given only half weight, Duran falls back toward the pack. Kody Clemens and Chisholm move near the top, while Bazzana remains around 1.48.

When defense is removed entirely, Kody Clemens becomes the leader and Bazzana rises to roughly third among the principal candidates.

Bazzana’s position is unusually stable. He does not need a favorable defensive estimate to make his case. Quite the opposite. Defense is holding his total down.

Duran’s candidacy is much more sensitive. His argument for being clearly superior to Bazzana rests heavily on trusting the defensive estimate.

That does not mean we should throw defense away. It means the gap between Duran and Bazzana is less certain than the raw 2.22 versus 1.43 WAR comparison initially suggests.

The Playing-Time Problem

There is another issue working against Bazzana. He simply had fewer opportunities.

Duran had 297 plate appearances. Chisholm had 334. Meidroth had 364. Clement had 342. Cole Young had 357.

Bazzana had 249.

That was not because he had been ineffective or injured for much of the major-league season. He did not make his debut until April 28.

One simple way to examine the effect of playing time is to normalize WAR to 600 plate appearances:

\mathrm{WAR}_{600} = 600 \left( \frac{ \mathrm{WAR} }{ \mathrm{PA} } \right)

For Bazzana:

\mathrm{WAR}_{600,\mathrm{Bazzana}} = 600 \left( \frac{ 1.432 }{ 249 } \right) \approx 3.45

That is not a projection. It does not mean Bazzana would necessarily have finished a full season with 3.45 WAR per 600 plate appearances.

It simply normalizes the rate at which value had accumulated. By that measure, Bazzana’s performance looks considerably more All-Star-like. His roughly 3.45 WAR-per-600 pace was competitive with the best players in the group.

The tradeoff is philosophical. Should an All-Star ballot reward the player who has accumulated the most value, or the player who has performed at the highest level while actually on the field?

There is no universally correct answer. If total contribution is the standard, Bazzana’s late arrival hurts him. If, on the other hand, quality of performance matters heavily, his case improves.

The Statcast Warning

There is, however, one piece of evidence that prevents me from becoming completely enthusiastic about Bazzana’s first-half numbers. His actual production was noticeably better than his Statcast expected production.

FanGraphs reported:

wOBA: .3319

xwOBA: .3025

SLG: .412

xSLG: .352

AVG: .250

xBA: .228

His barrel rate was only 4.2 percent, and his hard-hit rate was 36.1 percent.

We can express the difference between actual and expected wOBA as:

\Delta \mathrm{wOBA} = \mathrm{wOBA} - \mathrm{xwOBA}

For Bazzana:

\Delta \mathrm{wOBA}_{\mathrm{Bazzana}} = 0.3319 - 0.3025 = 0.0294

That is not trivial. It suggests that the quality of Bazzana’s batted-ball contact did not fully account for the results he produced.

But here again, context matters. Bazzana was not uniquely fortunate. A look at other players shows that Ernie Clement’s actual wOBA was about .324 while his xwOBA was only .272, an even larger gap. Duran also exceeded his xwOBA, although by a smaller amount, posting roughly .321 versus .303.

So Statcast weakens Bazzana’s case somewhat, but it does not destroy it. If anything, it warns us against treating the first-half offensive results of several second basemen as perfectly stable indicators of true talent. Figure 5 compares actual wOBA with Statcast expected wOBA for the second-base candidates. Players above the diagonal performed better than their contact quality would predict, while those below it underperformed their expected results. Bazzana is noticeably above the line.

The Duran Problem

At this point, one player becomes impossible to ignore, Ezequiel Duran.

If the question is simply, Who had accumulated the most FanGraphs value among the legitimate AL second-base candidates?, Duran has the cleanest answer.

He had 2.22 WAR, more plate appearances than Bazzana, and his offense was above average. And FanGraphs credited him with roughly eight defensive runs above average.
He was not an obscure candidate discovered only after the fact. MLB’s own All-Star voting updates showed Duran running near the top of the second-base voting while Bazzana was also among the leading candidates.

If I were selecting the reserve strictly from the FanGraphs value totals available at the time, I would choose Duran over Bazzana. But I would attach an asterisk to that conclusion. The size of Duran’s advantage depends heavily on defense, and half-season defensive values carry more uncertainty than half-season batting statistics. Take away the defensive separation, and Bazzana suddenly looks extremely competitive.

That makes the decision closer than the raw WAR totals imply.

And What About Clement?

There is also an interesting irony in the comparison with Ernie Clement. Clement was not merely elected to start. He became the leading American League vote-getter in Phase 1, automatically locking down the starting assignment at second base.

Statistically, however, Bazzana’s FanGraphs case was at least as strong.

Clement had 1.38 WAR to Bazzana’s 1.43. He had a 108 wRC+ to Bazzana’s 114. His baserunning contribution was negative while Bazzana’s was strongly positive. Both carried slightly negative FanGraphs defensive values in this snapshot.

Clement did have the playing-time advantage, 342 plate appearances compared with Bazzana’s 249. But if someone accepts Clement as a legitimate All-Star starter on performance grounds, it becomes difficult to argue that Bazzana was nowhere near All-Star caliber.

That comparison strengthens Bazzana’s defense considerably.

What Were the Players Seeing?

Statistics cannot tell us why players, managers, and coaches voted for Bazzana, but they allow us to construct a plausible explanation.

He was a rookie who had almost immediately become an above-average major-league hitter. He controlled the strike zone. He reached base. He ran aggressively and effectively. He was producing roughly 14 percent better offense than league average, and he had accumulated 1.4 WAR despite missing the opening month of the major-league season.

Players may also perceive certain skills more directly than statistical models do. They see bat speed, pitch recognition, baserunning pressure, defensive positioning, approach, and the quality of at-bats firsthand.

Reputation may have helped Bazzana. Being the first overall draft pick certainly made him familiar. But reputation alone is not a satisfactory explanation for the vote.

The numbers were already substantial enough to support it.

So, Did Travis Bazzana Deserve to Be an All-Star?

After reviewing the data, I would split the verdict into two parts. Did Travis Bazzana perform well enough to deserve serious All-Star consideration? Yes.

That conclusion is stronger than I expected when I began the analysis. A 114 wRC+, .341 OBP, positive baserunning value, and 1.43 WAR in only 249 plate appearances is an All-Star-caliber performance over the time he actually played.

His selection was not some inexplicable triumph of hype over production.

But the second question produces a different answer. Was Travis Bazzana clearly the most deserving reserve second baseman in the American League? No.

Ezequiel Duran had the strongest FanGraphs total-value case. Jazz Chisholm Jr., Chase Meidroth, and Kody Clemens also had legitimate statistical arguments. Bazzana’s relatively limited playing time and slightly negative defensive value keep him from being the obvious choice.

If I had been required to select one reserve strictly on the performance data available through July 4, I probably would have selected Duran.

Still, “probably Duran” is very different from “Bazzana did not deserve it.”

That distinction is the central finding of this exercise.

A Better Verdict

My original reaction to Bazzana’s selection was essentially: Really? Already?

The data changed that reaction. The best description of the selection is that it was not undeserved. Nor is it obvious but it is defensible.

Bazzana was already one of the strongest offensive second basemen in the American League. He added meaningful value on the bases. His WAR total was competitive despite substantially fewer plate appearances than most of the other leading candidates. Defense weakened his candidacy, and his Statcast-expected numbers suggest that some of his offensive production exceeded the quality of his contact.

Those are real reservations. But they are reservations about whether he was the best choice, not whether he belonged in the conversation.

That may be the most interesting thing about All-Star selections. The statistics often do not identify one incontrovertible winner. They reveal the tradeoffs.

Duran offered accumulated value and defense.

Chisholm offered defense and speed.

Kody Clemens offered offensive power.

Meidroth offered a balanced package with substantial defensive value.

Bazzana offered offense, on-base ability, speed, and an impressive rate of value accumulation in a shortened first half.

The Player Ballot chose Bazzana. The numbers say that choice was arguable yet reasonable.

My guess is that there will be many more All-Star games in Bazzana’s future. He is an impressive young player.

 

How Many Pitching Statistics Do We Really Need?

How Many Pitching Statistics Do We Really Need?

Correlation, Redundancy, and the Search for Independent Information in Baseball’s Pitching Metrics

Baseball has no shortage of pitching statistics.

ERA is still with us. So is WHIP. But they now share the stage with FIP, xFIP, SIERA, K-BB%, WAR, BABIP, ground-ball percentage, strikeout percentage, walk percentage, home-run rate, and an expanding collection of increasingly specialized measures.

That creates an interesting problem.

Are all these statistics actually telling us different things?

Or have we created many different ways of describing the same underlying pitching abilities?

The question became especially interesting after my study of WHIP. WHIP was very strongly associated with same-season ERA, but it was considerably weaker at predicting ERA one year later. FIP, xFIP, SIERA, and K-BB% all performed better as forward-looking measures.

That suggested a different question.

Perhaps we should stop asking which statistic is “best.”

Instead, we should ask:

How much independent information does each statistic actually contain?

That is the question I investigate here.

The Data

I used the same season-level FanGraphs dataset covering 2002 through 2025.

The data include season, innings pitched, ERA, FIP, xFIP, WAR, BABIP, home-run rate, and related pitching measures. A separate rate-statistics export includes WHIP, K%, BB%, K-BB%, ERA-, FIP-, xFIP-, FIP, xFIP, and SIERA. The batted-ball data add ground-ball percentage and related contact measures.

As in the WHIP study, I required at least 100 innings in a season:

IP_{i,y} \geq 100

That produced 1,699 qualifying pitcher-seasons.

For the predictive analysis, a pitcher had to reach 100 innings in consecutive seasons:

IP_{i,y} \geq 100 \quad \text{and} \quad IP_{i,y+1} \geq 100

That left 927 consecutive-season pairs involving 299 different pitchers.

Because no pitcher reached 100 innings during the shortened 2020 season, that year naturally drops out of the consecutive-season analysis.

The Metrics

The main study included:

ERA, WHIP, FIP, xFIP, SIERA, K%, BB%, K-BB%, HR/9, BABIP, GB%, and WAR.

Some of these variables are obviously related.

One relationship is exact:

\mathrm{K\!-\!BB\%} = K\% - BB\%

This simple equation illustrates the broader issue surprisingly well.

K%, BB%, and K-BB% may occupy three separate columns on a leaderboard, but they do not represent three independent pieces of information.

K-BB% is constructed directly from the other two.

The relationships among FIP, xFIP, SIERA, ERA, WHIP, and the underlying pitching rates are more complicated.

But the same basic problem remains.

Different names do not necessarily mean different information.

First Look: The Correlation Matrix

The natural place to begin is with Pearson correlation.

Rather than displaying the variables in an arbitrary order, I used hierarchical clustering to place statistics with similar correlation structures near one another.

Figure 1. Correlation Structure of Major Pitching Metrics

Several relationships immediately stand out.

Metrics Correlation
xFIP and SIERA 0.970
K% and K-BB% 0.939
FIP and WAR -0.885
FIP and xFIP 0.873
SIERA and K-BB% -0.858
FIP and SIERA 0.854
ERA and WHIP 0.811

The relationship between xFIP and SIERA is extraordinary.

Their correlation is approximately:

r_{\mathrm{xFIP},\mathrm{SIERA}} = 0.970

Squaring that correlation gives:

R^2 = (0.970)^2 \approx 0.941

So roughly 94 percent of their observed variation is shared in a simple linear sense.

That does not make xFIP and SIERA identical. They are calculated differently and emphasize somewhat different aspects of pitching.

But statistically, they move together to an extraordinary degree.

ERA and WHIP tell a similar, though less extreme, story. Their same-season correlation is approximately 0.81.

WHIP and ERA look like different statistics.

In practice, they often move together.

Correlation Is Only the Beginning

Pairwise correlation cannot tell us everything.

A statistic may have only moderate correlations with several individual variables while still being highly predictable from all of them collectively.

This matters enormously in regression analysis.

Suppose we try to predict one variable using all the others. If that variable can already be reconstructed very accurately, then adding it to a large regression may provide very little truly new information.

A standard way to investigate this is the Variance Inflation Factor:

\mathrm{VIF}_j = \frac{ 1 }{ 1-R_j^2 }

Here, R_j^2 measures how well predictor (j) can itself be predicted by the other predictors.

A VIF near 1 suggests relatively little redundancy.

As VIF increases, multicollinearity becomes more serious.

I excluded K-BB% from this particular calculation because K%, BB%, and K-BB% have an exact mathematical dependency.

The remaining results were striking.

Figure 2. Multicollinearity Among the Pitching Metrics

Metric VIF
FIP 109.0
WHIP 93.1
K% 47.5
HR/9 46.9
xFIP 44.1
BB% 32.1
BABIP 29.2
SIERA 28.2
ERA 5.8
GB% 3.9

These values are enormous.

But they should not be interpreted as evidence that FIP, WHIP, or SIERA are bad statistics.

That is not what VIF measures.

The result, instead, shows that combining all these variables into a single regression equation creates extreme redundancy.

Several predictors are trying to explain the same underlying variation.

That will matter shortly.

What Should We Predict?

A statistic can look extremely impressive when it is asked to explain something happening in the same season.

Prediction is harder.

So, as in the WHIP study, I used next-season ERA as the principal target.

The general idea is:

\widehat{\mathrm{ERA}}_{i,y+1} = \beta_0 + \sum_{j=1}^{p} \beta_j z_{i,y,j}

where the predictors come from year (y), while the target is ERA in year (y+1).

The predictors were standardized:

z_{i,y,j} = \frac{ x_{i,y,j} - \overline{x}_{j} }{ s_j }

Standardization allows coefficients from variables measured on very different numerical scales to be compared more sensibly.

More importantly, I did not simply fit the models and report their in-sample (R2).

I used leave-one-season-out cross-validation.

One season was withheld.

The model was trained using all the other seasons.

Then it had to predict the observations belonging to the season it had not seen.

The procedure was repeated across the available seasons.

That creates a much more demanding test.

Which Single Metric Predicts Best?

Before constructing complicated regression models, it makes sense to give each statistic a chance by itself.

Figure 3. Which Single Metric Best Predicts Next-Season ERA?

The results were:

Predictor in year (y) Predictive (R^2) for ERA in (y+1)
SIERA 0.206
FIP 0.199
xFIP 0.195
K-BB% 0.178
K% 0.176
WAR 0.145
ERA 0.116
WHIP 0.101
HR/9 0.059
GB% -0.009
BB% -0.009
BABIP -0.012

SIERA wins.

But only narrowly.

FIP and xFIP are very close behind it.

That is exactly what we might expect from the correlation matrix. If several statistics contain much of the same information, their predictive performance should often be similar.

K-BB% may be the most impressive result in the table.

It is remarkably simple:

\mathrm{K\!-\!BB\%} = K\% - BB\%

Yet its predictive (R2) reaches approximately 0.178.

That is not far behind FIP, xFIP, and SIERA.

BABIP performs particularly poorly. Its cross-validated (R2) is slightly negative.

A negative predictive (R2) does not mean that higher BABIP magically predicts lower ERA.

It means something simpler.

For this particular prediction problem, the fitted BABIP model performs slightly worse than simply predicting the average ERA.

What Happens If We Put Everything Into One Regression?

Here is where things become interesting.

I constructed a full ordinary least-squares regression containing:

ERA, WHIP, FIP, xFIP, SIERA, K%, BB%, HR/9, BABIP, and GB%.

K-BB% was omitted because including K%, BB%, and K-BB% together would introduce exact linear dependency.

The model then produced something strange.

The standardized coefficient for WHIP was approximately:

\beta_{\mathrm{WHIP}} \approx -0.389

Taken literally, that would imply that a higher WHIP predicts a lower future ERA, once the other statistics are held constant.

That is not a sensible baseball interpretation.

It is a multicollinearity problem.

Figure 4. OLS, Ridge, and LASSO Coefficients

Ordinary least squares tries to divide explanatory credit among variables that contain overlapping information.

That can make individual coefficients unstable.

One variable gets a large positive coefficient.

Another highly related variable gets a negative coefficient.

A small change in the sample can alter them again.

The overall model can still predict reasonably well.

The individual coefficients, however, become difficult to interpret.

This is exactly why simply adding every available baseball statistic to a regression is not necessarily a good idea.

More variables do not automatically produce more knowledge.

Sometimes they produce more confusion.

Ridge Regression

Ridge regression offers one solution.

Rather than allowing coefficients to become arbitrarily large, Ridge penalizes them:

\min_{\beta} \left[ \sum_{i=1}^{n} \left( y_i-\widehat{y}_i \right)^2 + \lambda \sum_{j=1}^{p} \beta_j^2 \right]

The tuning parameter (lambda) determines how strongly the coefficients are shrunk toward zero.

When the predictors contain large amounts of overlapping information, this can make the model considerably more stable.

That is exactly what happened.

Under ordinary least squares, the standardized WHIP coefficient was approximately -0.389.

Under Ridge regression, it became approximately:

\beta_{\mathrm{WHIP,Ridge}} \approx 0.013

Essentially zero.

That is an important result.

The model is not saying WHIP is useless.

It is saying:

Once all the other pitching information is already known, WHIP contributes very little additional information about next-season ERA.

That is a very different statement.

LASSO: Let the Model Throw Statistics Away

LASSO takes the regularization idea further.

Instead of penalizing squared coefficients, it penalizes their absolute values:

\min_{\beta} \left[ \sum_{i=1}^{n} \left( y_i-\widehat{y}_i \right)^2 + \lambda \sum_{j=1}^{p} \left| \beta_j \right| \right]

This has an interesting consequence.

LASSO can set some coefficients exactly to zero.

That turns our original question into an empirical experiment.

Give the model all these pitching statistics.

Then ask:

Which ones does it decide it does not need?

In the full-sample fit using the cross-validated penalty, LASSO retained four nonzero predictors:

FIP

SIERA

K%

HR/9

ERA went to zero.

WHIP went to zero.

xFIP went to zero.

BB%, BABIP, and GB% went to zero.

This does not mean those discarded statistics contain no useful baseball information.

It means that, for predicting next-season ERA after the retained variables were already available, their additional contribution was small enough that LASSO discarded them.

That is precisely the kind of redundancy we set out to investigate.

How Stable Was LASSO’s Decision?

A single LASSO fit is useful, but correlated predictors can substitute for one another.

So I also examined how often each statistic survived across the outer cross-validation folds.

Figure 5. Which Metrics Does LASSO Keep?

The approximate selection frequencies were:

Metric Selected
FIP 100%
SIERA 100%
K% 100%
HR/9 90%
ERA 48%
BB% 24%
xFIP 19%
WHIP 10%
BABIP 10%
GB% 5%

Three statistics survived every time:

FIP, SIERA, and K%.

HR/9 survived in approximately 90 percent of the folds.

WHIP survived only about 10 percent of the time.

That is particularly interesting after the previous WHIP study.

WHIP is very good at describing current run prevention.

Yet once the regression already knows FIP, SIERA, strikeout rate, home-run rate, and the other variables, WHIP rarely contains enough unique predictive information to survive LASSO.

That is not a contradiction.

It is the distinction between useful information and unique information.

Does the Giant Model Actually Predict Better?

Now we arrive at the question that matters most.

Perhaps all this redundancy does not matter if the large model predicts much better.

Does it?

Figure 6. More Metrics Help, but Only a Little

The cross-validated results were:

Model Predictive (R^2)
ERA only 0.116
SIERA only 0.206
Five-skill model 0.215
All metrics, OLS 0.213
All metrics, Ridge 0.222
All metrics, LASSO 0.217

The five-skill model used K%, BB%, HR/9, BABIP, and GB%.

The best model was the full Ridge regression.

Its predictive (R2) was approximately:

R^2_{\mathrm{Ridge}} = 0.222

SIERA alone produced:

R^2_{\mathrm{SIERA}} = 0.206

The improvement was therefore:

\Delta R^2 = 0.222 - 0.206 = 0.016

Just 1.6 percentage points.

The RMSE tells essentially the same story.

SIERA alone produced an RMSE of about 0.737 ERA runs.

The full Ridge model reduced that to approximately 0.729.

The larger model is better.

But only a little.

That may be one of the most important results in the study.

We gave the model a large collection of modern pitching statistics.

Most of the additional information barely moved the prediction.

Once We Know SIERA, What Else Helps?

This provides another way to look at redundancy.

Start with SIERA, the strongest individual predictor.

Then add other statistics one at a time.

Figure 7. Once We Know SIERA, Most Extra Metrics Add Very Little

SIERA alone:

R^2 = 0.206

Add FIP and the result improves to approximately:

R^2 = 0.217

That is a genuine, although modest, improvement.

Add xFIP to SIERA and predictive performance actually slips slightly, to approximately 0.204.

Add WHIP and it falls to roughly 0.203.

Adding K-BB% changes almost nothing.

Even combining SIERA, FIP, and xFIP reaches only about 0.216.

This is a remarkably clean demonstration of statistical redundancy.

Three statistics are not necessarily three times as informative as one.

Sometimes the second statistic is largely repeating the first.

The third repeats them both.

Why xFIP and SIERA Are a Good Example

Consider again:

r_{\mathrm{xFIP},\mathrm{SIERA}} = 0.970

If two statistics move almost perfectly together, there is simply not much room for one to add entirely new predictive information after the other is already known.

They can still differ conceptually.

Their formulas can still have different purposes.

Their disagreements can even be analytically useful.

But conceptual difference does not guarantee statistical independence.

That distinction is central to this study.

What About WAR?

ERA is only one possible definition of future pitching success.

So I repeated the regression analysis using next-season WAR as the target.

Current WAR itself was the strongest individual predictor:

R^2_{\mathrm{WAR}_y,\mathrm{WAR}_{y+1}} \approx 0.336

SIERA reached approximately 0.252.

FIP reached about 0.243.

xFIP was around 0.235.

WHIP was considerably weaker at approximately 0.125.

Figure 8. Predicting Next-Season WAR

The multivariable models produced:

Model Predictive (R2)
WAR only 0.336
WAR + IP 0.336
Skill model + IP 0.297
All metrics OLS 0.373
All metrics Ridge 0.375
All metrics LASSO 0.377

Here the large models add somewhat more information.

The best LASSO model improves predictive (R2) from roughly 0.336 using current WAR alone to approximately 0.377.

That is a more noticeable gain than we observed when predicting ERA.

Still, the central lesson remains.

Adding many statistics helps.

But the improvement is nowhere near proportional to the number of variables added.

The Strange WHIP Coefficient Revisited

The negative WHIP coefficient in the ordinary regression is worth returning to because it demonstrates an important statistical point.

WHIP by itself predicts future ERA in the expected direction.

Higher WHIP is associated with higher future ERA.

But once we tell an ordinary regression to hold ERA, FIP, xFIP, SIERA, K%, BB%, HR/9, BABIP, and GB% constant, the WHIP coefficient becomes negative.

Those are two very different questions.

The simple regression asks:

What happens to future ERA when WHIP changes?

The giant multiple regression asks:

What happens when WHIP changes while an enormous collection of closely related pitching statistics somehow remains fixed?

That second scenario may have very little resemblance to an actual pitcher.

The variables are too interconnected.

OLS nevertheless tries to divide the shared information among them.

The resulting coefficient is mathematically legitimate.

Its baseball interpretation is questionable.

Ridge responds by shrinking the WHIP coefficient almost to zero.

LASSO simply removes WHIP.

Both approaches give us a more sensible picture of what the data are saying.

How Many Statistics Do We Really Need?

There is no universal answer.

It depends on the question.

If I want to know how successfully a pitcher kept runners off base, WHIP is excellent.

If I want to know how many earned runs he actually allowed, ERA tells me exactly that.

If I want a simple forward-looking estimate of next-season ERA, SIERA performed best among the individual metrics examined here.

If I want to squeeze out every additional bit of predictive accuracy, a regularized multivariable model performs somewhat better.

But the key word is somewhat.

Going from SIERA alone to a ten-variable Ridge regression improved predictive (R2) from about 0.206 to 0.222.

The sophisticated model wins.

Barely.

The Bigger Lesson

A modern baseball leaderboard can create an illusion of enormous amounts of independent information.

Twenty columns look like twenty facts.

Statistically, that may not be true.

Strikeout ability appears directly in K%.

It appears again in K-BB%.

It also enters FIP, xFIP, and SIERA.

Walks do the same.

Home runs affect FIP and other estimators.

BABIP influences hit-based measures.

ERA and WHIP share the consequences of allowing baserunners.

The columns multiply faster than the underlying baseball phenomena.

That is not a criticism of advanced statistics.

Different metrics were created to answer different questions. They emphasize different aspects of performance. They make different assumptions. They may be useful in different contexts.

The mistake would be assuming that because two statistics have different names, they must contain completely different information.

They often do not.

Conclusion

The original question was simple:

How many pitching statistics do we really need?

The answer is more interesting than I expected.

Many pitching metrics are strongly correlated.

Some are extraordinarily correlated.

xFIP and SIERA correlate at approximately 0.97. ERA and WHIP correlate at about 0.81. FIP, xFIP, SIERA, strikeout measures, walk measures, and home-run measures overlap so heavily that placing all of them in the same ordinary regression produces severe multicollinearity.

The VIF analysis shows the problem.

The unstable OLS coefficients make it visible.

Ridge regression controls it.

LASSO begins throwing redundant statistics away.

And the predictive results show just how little we lose by simplifying.

SIERA alone explains about 20.6 percent of next-season ERA variation under leave-one-season-out prediction.

A full ten-variable Ridge model improves that to approximately 22.2 percent.

Ten statistics are better than one.

But not by much.

That may be the most important conclusion.

The goal should not be to collect the largest possible number of metrics.

It should be to identify statistics that represent genuinely different dimensions of pitching.

After that point, we increasingly begin measuring the same underlying abilities again.

And again.

Just under different names.

 

WHIP Is Simple. But Is It Actually Predictive?

WHIP may be one of baseball’s most intuitive pitching statistics.

It asks a straightforward question: how many runners does a pitcher allow to reach base by hit or walk per inning pitched?

\mathrm{WHIP}=\frac{BB+H}{IP}

Lower is better.

There is an appealing logic to the statistic. A pitcher who keeps runners off base should allow fewer runs. A pitcher who repeatedly permits hits and walks should eventually pay for them.

And, within a single season, that logic works remarkably well.

But prediction is a different problem.

The fact that WHIP describes a pitcher’s current performance does not necessarily mean that it tells us much about what he will do next year.

So I wanted to test a more difficult question:

How predictive is WHIP?

The Data

I used the season-level FanGraphs data covering 2002 through 2025. The exports include season, innings pitched, ERA, FIP, xFIP, WAR, BABIP, and related measures. A second export supplies the underlying hits and walks needed to examine the components of WHIP. The rate-statistics file contains WHIP itself along with K%, BB%, K-BB%, ERA-, FIP-, xFIP-, FIP, xFIP, and SIERA.

I imposed a minimum of 100 innings:

IP_{i,y}\geq100

That produced 1,699 qualifying pitcher-seasons.

For the year-to-year analysis, a pitcher had to reach 100 innings in consecutive seasons:

IP_{i,y}\geq100\quad\text{and}\quad IP_{i,y+1}\geq100

That left 927 consecutive-season pairs involving 299 different pitchers.

The shortened 2020 season effectively drops out under this rule. I also checked a less restrictive version allowing the following season to fall below 100 innings. The conclusions barely changed, so the principal findings are not an artifact of the cutoff.

First, Does WHIP Describe Current Performance?

Very much so.

The relationship between WHIP and same-season ERA is striking.

Figure 1. Same-Season WHIP and ERA

The fitted relationship is:

\mathrm{ERA}=-1.591+4.316(\mathrm{WHIP})

with:

R^2=0.657

The correlation is approximately:

r=0.811

So WHIP explains about 65.7 percent of the observed variation in same-season ERA among these pitcher-seasons.

That is enormous for such a simple statistic.

This should not be surprising. Runs require baserunners, and WHIP measures two of the principal ways pitchers put those runners on base.

But this is description, not prediction.

The real test begins when we move forward one year.

Does WHIP Repeat Itself?

Before asking whether WHIP predicts ERA, it makes sense to ask an even simpler question.

Does a pitcher’s WHIP this year predict his WHIP next year?

Figure 2. Year-to-Year WHIP

The regression is:

\mathrm{WHIP}_{y+1}=0.626+0.511(\mathrm{WHIP}_{y})

with:

R^2=0.224

The year-to-year correlation is approximately 0.474.

That is meaningful.

Pitchers with low WHIPs tend to have relatively low WHIPs again. Pitchers with high WHIPs tend to remain higher.

But look at what happened to the explanatory power.

Same season:

R2=0.657

Next season:

R2=0.224

More than two-thirds of WHIP’s apparent explanatory power disappears when we move the target forward one season.

That is our first indication of regression toward the mean.

The Main Question: Does WHIP Predict Next-Season ERA?

Now we come to the central test.

Figure 3. WHIP and Next-Season ERA

The fitted equation is:

\mathrm{ERA}_{y+1}=1.482+1.918(\mathrm{WHIP}_{y})

with:

R^2=0.110

The correlation is approximately 0.332.

So current WHIP explains only about 11 percent of the variation in next-season ERA.

That is not nothing.

WHIP contains a real signal.

But compare 11 percent with the 66 percent relationship between WHIP and ERA in the same season.

That difference is the heart of the study.

WHIP is much better at telling us what has happened than telling us what will happen.

How Does WHIP Compare With Other Predictors?

This is where the result becomes more interesting.

I compared several current-season statistics as predictors of next-season ERA:

Current-season statistic Leave-one-season-out predictive (R2)
SIERA 0.206
FIP 0.199
xFIP 0.195
K-BB% 0.178
ERA 0.116
WHIP 0.101
BABIP -0.012

For this comparison, I used leave-one-season-out validation. Each season was withheld in turn, the regression was estimated using the other seasons, and predictions were generated for the season that had not been used to fit the model.

That makes this a considerably tougher test than simply reporting an in-sample correlation.

Figure 4. Which Statistic Predicts Next-Season ERA Best?

SIERA wins.

FIP and xFIP are essentially right behind it.

K-BB% performs surprisingly well considering how simple it is.

Current ERA beats WHIP, although not by much.

And BABIP is essentially useless for predicting next-season ERA in this sample. Its cross-validated (R^2) is actually slightly below zero, meaning that simply predicting the overall mean would perform marginally better.

This ranking makes baseball sense.

FIP, xFIP, and SIERA deliberately concentrate more heavily on pitcher-controlled characteristics such as strikeouts, walks, and home runs.

WHIP includes hits.

That is both its strength and its weakness.

Hits matter enormously right now.

They are less stable going forward.

Regression Toward the Mean

One of the clearest ways to see the problem is to divide the pitchers into five groups based on their WHIP in year (y).

Figure 5. WHIP Quintiles and Regression Toward the Mean

The best-WHIP quintile averaged approximately:

1.054 WHIP and 2.97 ERA in the current season.

Their average ERA the following year?

3.47.

That is still very good, but nowhere near 2.97.

At the other extreme, the worst-WHIP quintile averaged approximately:

1.454 WHIP and 4.69 ERA.

The following season their ERA improved to approximately:

4.22.

The extremes move toward one another.

The spread between the best and worst groups was about 1.72 ERA runs in year (y).

One season later, it was only about 0.76 runs.

WHIP clearly contains information that persists.

It simply does not persist nearly as strongly as the current-season relationship might lead us to believe.

Which Part of WHIP Is Responsible?

WHIP consists of two pieces:

\mathrm{WHIP}=\frac{H}{IP}+\frac{BB}{IP}

That gives us a useful experiment.

We can separate hits from walks.

Using per-nine-inning versions:

H/9=9\left(\frac{H}{IP}\right)

and:

BB/9=9\left(\frac{BB}{IP}\right)

The results contain a small surprise.

Figure 6. What Part of WHIP Persists?

BB/9 is actually the most stable individual component from one year to the next.

Its year-to-year (R2) is approximately:

0.425

H/9 has a year-to-year (R2) of about:

0.257

and WHIP itself:

0.224

So walk rate is considerably more persistent than hit rate.

But when the target is next-season ERA, the story reverses.

Current H/9 explains about 10.4 percent of future ERA variation.

Current BB/9 explains only about 1.1 percent.

WHIP explains about 11.0 percent.

That sounds contradictory at first.

It is not.

Walk rate is a repeatable pitcher skill, but within this group of durable pitchers its variation alone does not explain very much of next year’s ERA. Hits allowed are less stable, yet they contain information related to strikeout ability, contact suppression, and other pitcher characteristics.

Importantly, BABIP itself had essentially no relationship with next-season ERA.

So the useful information in H/9 is not simply “this pitcher had a low BABIP, therefore he will again.”

Does WHIP Add Anything to Better Pitching Metrics?

This may be the most revealing test of all.

Suppose we already know a pitcher’s FIP or SIERA.

Does knowing his WHIP improve our prediction?

Figure 7. Does WHIP Add Predictive Information?

The answer is essentially no.

The leave-one-season-out results were approximately:

Model Predictive (R2)
ERA 0.116
ERA + WHIP 0.119
FIP 0.199
FIP + WHIP 0.197
SIERA 0.206
SIERA + WHIP 0.203

Adding WHIP to ERA produces a tiny improvement.

Adding WHIP to FIP actually lowers out-of-sample performance slightly.

Adding it to SIERA does the same.

That does not mean WHIP is worthless.

It means that once we know a more sophisticated pitcher estimator, WHIP contributes little additional predictive information.

SIERA already knows much of what we need to know about the pitcher’s underlying skills.

WHIP mainly tells us what happened to those skills and batted balls in the season we just observed.

A League-Adjusted Check

There is another potential concern.

Baseball’s run environment changed between 2002 and 2025.

Fortunately, the supplied FanGraphs data also contain ERA-, FIP-, xFIP-, and league-adjusted WHIP+.

Repeating the year-to-year analysis with those adjusted measures does not change the conclusion.

Current WHIP+ explains about 9.4 percent of next-season ERA- variation.

Current ERA- explains about 10.1 percent.

FIP- rises to roughly 17.6 percent, while xFIP- reaches approximately 17.9 percent.

The same hierarchy remains.

The finding is therefore not merely the result of changing league scoring environments.

What WHIP Is Good For

None of this makes WHIP a bad statistic.

Quite the opposite.

A statistic explaining roughly 66 percent of the variation in same-season ERA with a formula consisting only of hits, walks, and innings is extraordinarily efficient.

WHIP answers a very useful question:

How successfully has this pitcher kept runners off base?

It is easy to calculate.

Easy to understand.

Easy to compare.

And strongly connected to what happened on the scoreboard.

But a descriptive statistic and a predictive statistic are not the same thing.

That distinction matters.

FIP, xFIP, SIERA, and K-BB% sacrifice some of WHIP’s intuitive simplicity in an attempt to isolate skills that are more likely to persist.

Our results suggest that tradeoff works.

Conclusion

WHIP passes one test spectacularly and another only modestly.

As a description of current performance, it is excellent.

R^2_{\mathrm{WHIP,\ same\ season\ ERA}}=0.657

As a predictor of its own next-season value:

R^2_{\mathrm{WHIP}_y,\mathrm{WHIP}_{y+1}}=0.224

And as a predictor of next-season ERA:

R^2_{\mathrm{WHIP}_y,\mathrm{ERA}_{y+1}}=0.110

Under the stricter leave-one-season-out test, that last value falls slightly to about 0.101.

SIERA roughly doubles that predictive performance.

FIP and xFIP come close.

K-BB% also comfortably beats WHIP.

So the answer to the original question is fairly clear.

WHIP is predictive, but not especially predictive.

A low WHIP today is meaningful evidence that a pitcher will perform well tomorrow. It is simply much weaker evidence than its strong relationship with today’s ERA might suggest.

That may be WHIP’s most interesting characteristic.

It looks like a forward-looking statistic because it is so closely connected with pitching success.

In reality, much of its power belongs to the present.

WHIP tells us exceptionally well what a pitcher has done.

For predicting what he will do next, we can do better.

 

Third Base as Baseball’s Hybrid Position

Introduction

This project began with a simple question:

Who were the greatest third basemen?

But the longer the study went on, the less simple that question became.

Third base is not first base. It is not a position where offense alone usually defines the job. But it is also not shortstop or second base, where defensive value and up-the-middle responsibility have historically borne the greater burden.

Third base lives between those worlds.

It is a reaction position.
It is an arm-strength position.
It is a power position.
It is a two-way position.

That is why third base is so interesting analytically. The position resists one-dimensional ranking. A player can dominate the position through offense, as Mike Schmidt, Eddie Mathews, Chipper Jones, George Brett, and Jose Ramirez did. A player can dominate it through defense, as Brooks Robinson did. A player can become historically important through balance, as Scott Rolen, Adrian Beltre, Nolan Arenado, and Wade Boggs did in different ways.

The point of this final chapter is not to introduce another ranking. It is to bring the project together.

The central claim is this:

Third base is baseball’s hybrid position.

It is not merely a corner power position.
It is not merely a defensive position.
It is the place where power, reaction defense, arm strength, durability, and era context all meet.

A Note on Era Context

Comparing baseball players across eras is difficult. The game has changed. The league has expanded. The talent pool has changed. Integration, globalization, training, travel, medicine, equipment, ballparks, strategy, and offensive environments all matter.

The Full House Modeling paper by Yan, Burgos, Kinson, and Eck makes this point directly. Their framework builds on Stephen Jay Gould’s “full house” idea by treating player achievement as part of a broader distribution of performance, while also considering the changing baseball talent pool over time. Their method is more ambitious than the one used here because it explicitly models talent-pool size and era-adjusted performance.

This project does not implement Full House Modeling. It uses a narrower, more transparent adjustment:

Compare third basemen to other third basemen in the same season.

That is the role of the z-score method.

z_{i,y,m} = \frac{ x_{i,y,m} - \overline{x}_{p,y,m} }{ s_{p,y,m} }

Where:

i = \text{player} y = \text{season} p = \text{primary position} m = \text{metric}

In plain English, the question becomes:

How far did this third baseman stand above or below other third basemen in the same season?

That does not solve every era problem. But it does control for one of the most important things: the player’s immediate positional environment.

Defining the Position

Before ranking third basemen, we first need to understand the position itself.

For the home run study, each player-season was assigned a primary position based on where the player appeared most often:

\mathrm{PrimaryPosition}_{i,y} = \operatorname*{arg\,max}_{p \in \mathcal{P}} G_{i,y,p}

Where:

G_{i,y,p} = \text{games played by player } i \text{ at position } p \text{ in season } y

The main home run study used regular player-seasons:

PA \geq 300

The full time frame was:

1871–2025

That gave us a broad historical view of how home runs have been distributed by position.

Figure 1: Total Home Runs by Position

Figure 1. Total home runs by primary position, regular player-seasons, PA ≥ 300, 1871–2025.

The total home run bar chart gives the project a useful starting point. It shows where the historical home run volume has come from.

For regular player-seasons:

\mathrm{TotalHR}_{1B} = 48{,}064 \mathrm{TotalHR}_{RF} = 42{,}977 \mathrm{TotalHR}_{LF} = 40{,}663 \mathrm{TotalHR}_{3B} = 36{,}127

Third base ranks behind first base and the corner outfield positions, but ahead of center field, catcher, second base, and shortstop.

That matters.

It shows that third base has historically been a significant power source. It is not a middle-infield power-light position. But it is not first base either. Third base carries a heavier defensive burden than first base and a higher power expectation than second base or shortstop.

That is the hybrid identity of the position.

The total home run formula is simple:

\mathrm{TotalHR}_{p} = \sum_{i,y} HR_{i,y,p}

The conclusion is not subtle:

Third base belongs among the power-producing positions,

but it is not purely a power position.

Figure 2: Home Run Distribution by Position

Figure 2. Home runs by primary position, regular player-seasons, PA ≥ 300, 1871–2025.

The box plots add another layer.

The bar chart showed total historical volume. The box plot shows the distribution of individual player-seasons.

For regular player-seasons, the median home runs by position were:

DH: 18

1B: 12

RF: 11

LF: 10

3B: 9

C: 7

CF: 7

2B: 5

SS: 4

In compact form:

\widetilde{\mathrm{HR}}_{\mathrm{DH}} > \widetilde{\mathrm{HR}}_{\mathrm{1B}} > \widetilde{\mathrm{HR}}_{\mathrm{RF}} > \widetilde{\mathrm{HR}}_{\mathrm{LF}} > \widetilde{\mathrm{HR}}_{\mathrm{3B}} > \widetilde{\mathrm{HR}}_{\mathrm{C}} \approx \widetilde{\mathrm{HR}}_{\mathrm{CF}} > \widetilde{\mathrm{HR}}_{\mathrm{2B}} > \widetilde{\mathrm{HR}}_{\mathrm{SS}}

Again, third base sits in the middle-to-upper part of the power spectrum.

That position explains why third-base rankings are hard. A third baseman who hits like a first baseman is unusually valuable. A third baseman who fields like an elite shortstop-adjacent defender is also unusually valuable. A third baseman who does both is historically rare.

The Offensive Summit

The offensive studies showed that Mike Schmidt remains the defining offensive third baseman.

Using the Model C offensive framework, the season score combined on-base ability, isolated power, walk rate, strikeout avoidance, baserunning, run scoring, and run production.

\begin{aligned} \mathrm{ModelCOffensiveScore} &= z_{\mathrm{OBP}} + z_{\mathrm{ISO}} + z_{\mathrm{BB/PA}} + z_{\mathrm{LowSO/PA}} \\ &\quad+ z_{\mathrm{NetSB/PA}} + z_{\mathrm{R/PA}} + z_{\mathrm{RBI/PA}} \end{aligned}

The career score was the sum of weighted season scores:

\mathrm{CareerScore} = \sum_s \mathrm{WeightedSeasonScore}_{s}

And peak value was measured through a player’s best seven seasons:

\mathrm{Peak7Score} = \sum_{k=1}^{7} \mathrm{BestSeasonScore}_{k}

Schmidt’s greatness came from both career value and peak value. He was not merely a compiler. He was not merely a short-peak slugger. He combined long-term production with enormous seasonal dominance.

The offensive group around him included Eddie Mathews, Chipper Jones, George Brett, Jose Ramirez, Alex Rodriguez, Wade Boggs, Home Run Baker, Ron Santo, and David Wright.

But they were not all the same kind of offensive player.

Schmidt was power and patience.
Boggs was on-base dominance.
Brett was batting skill and all-around production.
Chipper was switch-hitting offensive force.
Jose Ramirez was compact modern power, baserunning, and peak efficiency.
Mathews was long-career left-handed power.

The offensive rankings showed that third-base greatness has several offensive forms.

The Defensive Summit

The traditional defensive study told a different story.

The defensive model used assists per game, putouts per game, double plays per game, fielding percentage, and low errors per game.

\mathrm{TraditionalDefensiveScore} = z_{\mathrm{A/G}} + z_{\mathrm{PO/G}} + z_{\mathrm{DP/G}} + z_{\mathrm{FPct}} + z_{\mathrm{LowE/G}}

The defensive peak and career results were clear:

Brooks Robinson is the traditional defensive summit.

Robinson stood apart because his defensive career was both excellent and enormous. He was not only a peak defender. He was a long-duration defensive institution.

The next group included Nolan Arenado, Gary Gaetti, Willie Kamm, Buddy Bell, Clete Boyer, Adrian Beltre, Scott Rolen, Matt Chapman, Graig Nettles, Mike Lowell, and Wade Boggs.

This defensive list is important because it changes the shape of the third-base conversation.

If we looked only at offense, Brooks Robinson would not appear near the top. If we looked only at defense, Chipper Jones would not appear near the top. Third base requires us to keep both dimensions visible.

Figure 3: Offense and Defense Together

Figure 3. Career offensive score versus traditional defensive score for third basemen.

The two-dimensional scatterplot is one of the most important figures in the project.

It shows that third basemen do not fall along a single line from bad to great. They occupy different regions of the offensive-defensive space.

The combined two-way score was built by standardizing the career offensive and defensive dimensions:

\mathrm{CombinedZ}_{i} = z_{\mathrm{Offense},i} + z_{\mathrm{Defense},i}

The top of the combined list included:

Mike Schmidt

Brooks Robinson

Scott Rolen

Nolan Arenado

Wade Boggs

Eddie Mathews

Chipper Jones

Willie Kamm

George Brett

Adrian Beltre

This is the key result of the project.

Mike Schmidt wins because he was historically great on offense and still positive on defense.

Brooks Robinson ranks so high because his defensive value was so extreme that it overcame a more modest offensive profile.

Scott Rolen and Adrian Beltre matter because they represent the two-way ideal. They were not Schmidt-level offensive forces or Brooks-level defensive singularities, but they were excellent in both directions.

Nolan Arenado appears as a modern defense-heavy two-way star. Wade Boggs appears as a batting-and-defense star whose value is not captured by power alone.

Third base becomes clearest when seen as a two-axis position.

Figure 4: Dendrogram of Top Combined Third Basemen

Figure 4. Dendrogram showing similarity among the top 15 third basemen by combined offense-defense score.

The combined top list is not the same as the top offensive or top defensive lists.

That is the point.

A one-stat answer hides the different paths to value. The combined score helps reveal those paths.

The general model is:

\mathrm{TwoWayScore}_{i} = z_{\mathrm{Offense},i} + z_{\mathrm{Defense},i}

But the interpretation matters more than the arithmetic.

Schmidt is the best answer because he dominates the total landscape.
Robinson is the best defensive answer.
Chipper is one of the best offensive answers.
Rolen and Beltre are among the cleanest two-way answers.
Boggs and Brett show that contact, on-base skill, and complete offensive value can compete with pure home run power.
Arenado shows how a modern elite defender can push into the historical top tier.

The final ranking should therefore not be read as a single flat list. It should be read as a map of greatness.

The Center of the Position

The middle of the project may have been just as revealing as the top.

When we searched for the most average third basemen, we found different kinds of centers.

Offensively, Casey Blake represented a kind of Model C center.
Defensively, Eddie Mathews appeared surprisingly close to the traditional defensive center.
In the two-dimensional offense-defense space, Bill Melton emerged as the closest player to the combined center.

The two-dimensional typicality calculation was:

\mathrm{Typicality}_{i} = \sqrt{ z_{\mathrm{Offense},i}^{2} + z_{\mathrm{Defense},i}^{2} }

A low value means the player is close to the center of the third-base population.

This was useful because the average player gives meaning to the extremes.

Without the center, Schmidt is just a big number.
With the center, Schmidt becomes distance from normal.
Without the center, Brooks Robinson is just a defensive legend.
With the center, we can see how unusual his profile really was.

The middle defines the scale.

Figure 5: Closest Third Basemen to the Offense-Defense Center

Figure 5. Third basemen closest to the two-dimensional offense-defense center, measured by typicality.

The Lower Tail

The lower-tail studies also mattered.

Ken Reitz appeared as the lower offensive tail among third-base regulars, but his traditional defensive profile was positive. That makes him more interesting than a simple “worst hitter” label.

Chris Johnson emerged as a lower-tier tailback because he combined a weak offensive standing with a weak traditional defensive standing.

Figure 6: Lowest Combined Offense-Defense Third Basemen

Figure 6. Lowest combined offense-defense scores among third-base regulars.

The lower-tail combined score was:

\mathrm{LowerTailScore}_{i} = z_{\mathrm{Offense},i} + z_{\mathrm{Defense},i}

The lowest combined players included:

Chris Johnson

Butch Hobson

Wes Helms

Larry Parrish

Harry Lord

Tom Brookens

Maikel Franco

Dean Palmer

Charley Smith

Ray Jablonski

The lower tail reinforced a point that runs through the whole project:

A bad offensive third baseman is not necessarily a bad third baseman.

A bad defensive third baseman is not necessarily a bad third baseman.

But a player weak in both dimensions falls quickly.

Third base gives players several ways to survive. But it also exposes players who cannot contribute either way.

Validation with WAR

The WAR validation study gave the project an external check.

The model was not designed to reproduce WAR exactly. But if the offensive and defensive scores were meaningful, they should explain a substantial portion of WAR variation.

At the career level for regular third basemen, the offense-plus-defense model was strong:

\mathrm{WAR} = 14.90 + 0.60(\mathrm{CareerOffensiveScore}) + 0.59(\mathrm{CareerDefensiveScore}) R^2 = 0.814

That is an important result.

It means the simple, transparent framework captures much of what WAR captures, even though WAR is built on a more complex run-value framework.

The validation did not prove the model was perfect. It showed that the model was sensible.

Offense mattered.
Defense mattered.
Together they explained much more than either alone.

Validation with wRC+

The wRC+ validation was equally useful because it directly tested the offensive side.

At the season level, Model C offensive score predicted wRC+ well:

wRC^+ = 101.47 + 5.86(\mathrm{ModelCOffensiveScore}) R^2 = 0.692

At the career level, the offensive score also connected strongly to career wRC+:

wRC^+ = 100.89 + 5.41(\mathrm{AverageOffensiveScore}) R^2 = 0.740

The defensive negative control was also important. Traditional defensive score did not meaningfully predict wRC+. That is what we wanted to see.

The offensive model tracked offense.
The defensive model tracked defense.
The two-way model tracked overall value.

That separation gives the project credibility.

Figure 7: WAR Validation

Figure 7. Career WAR, actual versus predicted from offensive and defensive scores.

The WAR validation figure belongs in the final chapter because it shows that the project is not only descriptive. It is also diagnostic.

The model found recognizable third-base greatness because its components corresponded to real baseball value.

Schmidt, Robinson, Beltre, Rolen, Boggs, Mathews, Brett, Chipper, Arenado, Santo, and others were not artifacts of the scoring system. They appeared because the scoring system captured meaningful dimensions of the position.

Figure 8: wRC+ Validation

Figure 8. Model C offensive score versus wRC+ for third-base seasons.

The wRC+ validation figure should follow the WAR figure.

Together, these two validation figures show the difference between offensive value and total value.

wRC+ confirms the offensive model.
WAR confirms the broader two-way model.

That distinction is exactly what third base requires.

The Dendrogram: Types of Greatness

The dendrogram in Figure 4 adds one final interpretive layer.

Instead of asking who ranked first, it asked which great players were similar.

The distance formula was:

d(i,j) = \sqrt{ \left( z_{\mathrm{Offense},i} - z_{\mathrm{Offense},j} \right)^2 + \left( z_{\mathrm{Defense},i} - z_{\mathrm{Defense},j} \right)^2 }

This produced recognizable clusters.

Schmidt, Chipper, Mathews, Brett, and Jose Ramirez formed an offense-heavy region.

Brooks Robinson stood apart as the defensive summit.

Rolen, Beltre, Arenado, Boggs, Nettles, Kamm, and similar players occupied the two-way or defense-strong region.

The dendrogram reinforces the project’s central point:

There is no single shape of third-base greatness.

There are families of greatness.

As Figure 4 shows, the dendrogram turns the ranking into a typology.

It lets the reader see why direct comparisons can be difficult.

Is Chipper Jones greater than Brooks Robinson?
It depends on what kind of greatness we are asking about.

Is Scott Rolen closer to Adrian Beltre than to Mike Schmidt?
In a two-dimensional offensive-defensive space, yes.

Is Mike Schmidt the most complete answer?
The evidence says yes.

Final Interpretation

After all of the rankings, models, box plots, bar charts, validations, residuals, and dendrograms, the final interpretation is this:

Mike Schmidt is the best overall third baseman in this framework.

Brooks Robinson is the defensive summit.

Chipper Jones is one of the offensive summits.

Scott Rolen and Adrian Beltre are two-way ideals.

Nolan Arenado is the modern defensive-power bridge.

Wade Boggs and George Brett show that third-base offense is not only home runs.

But the broader conclusion is about the position itself.

Third base is a test of balance.

A great third baseman can win with the bat.
A great third baseman can win with the glove.
The greatest third basemen usually do both.

The home run study showed that third base belongs in the power conversation. The defensive study showed that the position still carries serious fielding responsibility. The WAR and wRC+ validations showed that the model behaves sensibly. The center and lower-tail studies showed that the full distribution matters, not only the top.

That brings the project back to the “full house” idea. Baseball greatness is not only about the record holder. It is about where that record, season, or career sits within the full distribution of the game.

This project used third base as the test case.

And third base rewarded the approach.

Conclusion

Third base is baseball’s hybrid position.

It is close enough to first base and the corner outfield to demand power.
It is close enough to shortstop and second base to demand defense.
It is far enough from both extremes to create unusual player types.

That is why the position has produced so many different kinds of great players.

Mike Schmidt towers because he solved the hybrid problem better than anyone else in this framework. He hit like a historic slugger while remaining a positive defensive player at a demanding position.

Brooks Robinson towers because his defense was so exceptional that it reshaped the positional landscape.

Rolen, Beltre, Arenado, Boggs, Brett, Chipper, Mathews, Santo, and others fill out the map because each represents a different solution to the same positional problem.

The final answer, then, is not only a ranking.

It is a definition.

Third base is where power and defense meet.

And that is why it was worth studying.

 

 

The Evolution of the Baseball Player: Discovering Offensive Archetypes, 1954-2025

Baseball has changed enormously since the middle of the twentieth century.

Strikeouts have increased. Home-run rates have risen and fallen. Stolen bases have moved in and out of fashion. Relief pitching has become more specialized. Defensive positions have acquired different offensive expectations.

It is tempting, therefore, to divide baseball history into a sequence of distinct player types. The contact hitter belongs to one era. The base stealer belongs to another. The modern game belongs to the power hitter who walks frequently and strikes out even more frequently.

But is that really what happened?

Did one type of hitter replace another, or did the same basic offensive archetypes persist across changing statistical environments?

To explore that question, I used the Lahman baseball database to examine 9,218 qualified non-pitcher seasons from 1954 through 2025. I standardized each player against other qualified hitters in his own season, reduced the statistical profiles using principal component analysis, and then used clustering to identify recurring offensive archetypes.

The results reveal six recognizable types of offensive players:

  1. Patient-contact hitters
  2. Low-impact contact hitters
  3. Elite power-and-patience hitters
  4. Power-and-strikeout hitters
  5. Aggressive free-swingers
  6. Speed-and-contact hitters

Perhaps most importantly, the results suggest that baseball’s statistical environment has changed much more dramatically than its underlying distribution of player types.

The modern hitter may look different in the raw statistics. Relative to his contemporaries, however, he often occupies a role that has existed for generations.

Building the Historical Sample

I began with five Lahman tables:

  • Batting
  • People
  • Appearances
  • Fielding
  • Teams

Players who appeared for multiple teams during one season were combined into a single player-season record. I used the appearances data to assign each player a primary defensive position, defined as the position at which he appeared most frequently during that season.

Pitchers were excluded.

To account for seasons of different lengths, I defined a qualified season using the familiar standard of 3.1 plate appearances per scheduled team game.

Because teams occasionally played slightly different numbers of games, I used the median number of team games during each season:

PA_{i,y} \geq 3.1 \widetilde{G}_{y}

where:

\widetilde{G}_{y} = \operatorname{median}\left(G_y\right)

Here, (\widetilde{G}_{y}) represents the median number of team games played in season (y).

This method adjusts the qualification threshold for 154-game seasons, 162-game seasons, strike-shortened seasons, and the 60-game 2020 season.

The study begins in 1954 because the variables needed for the full model are not consistently complete before that date. Strikeouts become sufficiently complete before then, but caught stealing and sacrifice flies create additional limitations. Beginning in 1954 allows the same five-variable model to be used across the entire study period.

Figure 1. The number of qualified player-seasons generally increased as Major League Baseball expanded.

The increase in qualified seasons primarily reflects league expansion. The early portion of the study contains fewer than 100 qualified hitters in many seasons. By the late 1990s and early 2000s, the total frequently exceeded 150.

Measuring an Offensive Profile

I wanted to describe how a hitter produced offense, not merely how much offense he produced.

I therefore selected five variables:

  • On-base percentage
  • Isolated power
  • Walks per plate appearance
  • Strikeouts per plate appearance
  • Net stolen bases per plate appearance

Plate appearances were calculated as:

PA = AB + BB + HBP + SF + SH

On-base percentage was calculated as:

\mathrm{OBP} = \frac{ H + BB + HBP }{ AB + BB + HBP + SF }

Isolated power measures extra-base power beyond batting average:

\mathrm{ISO} = \mathrm{SLG} - \mathrm{AVG}

The walk and strikeout rates were:

\mathrm{BB/PA} = \frac{BB}{PA} \mathrm{SO/PA} = \frac{SO}{PA}

Finally, I defined net stolen-base production as:

\mathrm{NetSB/PA} = \frac{ SB - CS }{ PA }

This final measure rewards successful steals while penalizing caught-stealing events.

I did not include runs or RBI in the clustering model. Both statistics are strongly affected by batting order, teammates, and opportunity. They tell us something about the results of a player’s season, but less about the underlying style with which he produced those results.

Home runs were also not included as a separate rate because ISO already captures power production. Adding both ISO and home runs per plate appearance would have given power disproportionate weight in the clustering.

Baseball’s Changing Offensive Environment

The raw statistics immediately demonstrate how much the offensive environment changed.

Figure 2. Mean offensive rates among qualified hitters, 1954-2025.

In 1954, the average qualified hitter in the sample had approximately:

\mathrm{OBP} = 0.355 \mathrm{ISO} = 0.151 \mathrm{SO/PA} = 0.088

By 2025, the corresponding values were approximately:

\mathrm{OBP} = 0.330 \mathrm{ISO} = 0.178 \mathrm{SO/PA} = 0.204

The average qualified hitter’s strikeout rate more than doubled.

Power generally increased, particularly during the offensive surge of the late 1990s and again during the home-run-heavy seasons of the late 2010s. Walk rates changed much less. On-base percentage remained within a fairly narrow historical range, although it rose noticeably around 2000 before declining.

Stolen-base production followed a different pattern. It increased during the 1970s and 1980s, declined during the power-oriented environment that followed, and has recently begun to rise again.

These changes create a serious comparison problem. A 20 percent strikeout rate would have been extraordinary during much of the twentieth century. In the modern game, it may be close to ordinary.

A player cannot be classified historically based solely on raw statistics.

Adjusting Every Player for His Era

To make the seasons comparable, I standardized each variable within its own season.

For player (i), season (y), and offensive measure (m):

z_{i,y,m} = \frac{ x_{i,y,m} - \overline{x}_{y,m} }{ s_{y,m} }

A value of: z=0 indicates that the player was equal to the seasonal average.

A value of: z=1 indicates that he was one standard deviation above the seasonal average.

This adjustment alters the analysis’s meaning. I am not asking whether a hitter had a high strikeout rate in absolute terms. I am asking whether he struck out frequently compared with the other qualified hitters of his own season.

A player from 1965 and a player from 2025 can therefore belong to the same archetype even though their raw statistics differ substantially.

They occupied the same relative position within their respective baseball environments.

Reducing the Offensive Dimensions

The five standardized variables remain related to one another. High-OBP players often walk frequently. Power hitters may also strike out frequently. Speed-oriented players tend to have different power and contact profiles.

I used principal-component analysis to summarize these relationships.

The first principal component explained: 42.5% of the total variation.

The second explained: 25.0%. Together, the first two components explained: 42.5% +25.0% = 67.5%

The first three components explained approximately: 86.2% of the total variation.

The first component primarily represents overall offensive force. OBP, ISO, and walk rate all load positively on this dimension. Players far to the right of the PCA plot tend to reach base, hit for power, and draw walks at rates well above their seasonal environments.

The second component separates high-power, high-strikeout hitters from lower-strikeout players with stronger contact or speed characteristics.

How Many Archetypes Are There?

Clustering always requires a choice about how much detail to preserve.

A model with too few clusters combines meaningfully different players. A model with too many clusters produces distinctions that may be statistically fragile or difficult to interpret.

The k-means procedure attempts to minimize the total squared distance between each player and the center of his assigned cluster:

\mathrm{WCSS} = \sum_{k=1}^{K} \sum_{i \in C_k} \left\lVert \mathbf{z}_i - \boldsymbol{\mu}_k \right\rVert^2

where:

C_k=\mathrm{cluster}\ k

\mathbf{z}_i = \mathrm{standardized\ profile\ of\ player\!-\!season}\ i

\boldsymbol{\mu}_{k} = \mathrm{center\ of\ cluster}\ k

I tested solutions containing four through nine clusters.

Figure 3. The silhouette score is highest for the four-cluster model, although the six-cluster model retains useful baseball distinctions.

The silhouette statistic compares the average distance between a player and his own cluster with the distance between that player and the nearest alternative cluster:

s(i) = \frac{ b(i)-a(i) }{ \max\left\{a(i),b(i)\right\} }

The four-cluster solution produced the strongest formal separation. However, it merged several historically meaningful offensive styles.

In particular, it tended to combine elite power-and-patience hitters with less complete power hitters, and it reduced distinctions between contact-oriented players.

I therefore selected the six-cluster solution.

This is an interpretive decision. The six clusters are not six perfectly isolated biological species. Baseball players exist along continuous statistical dimensions. The purpose of the clusters is to provide a useful map of that continuum.

Figure 4. The 9,218 player-seasons displayed in the space formed by the first two principal components.

The overlap in Figure 4 is important. The clusters are recognizable, but their boundaries are not absolute. A player near a boundary may resemble members of two neighboring archetypes.

The Six Offensive Archetypes

Figure 5 shows the standardized center of each cluster.

Figure 5. Each value represents the number of standard deviations above or below the seasonal mean.

1. Patient-Contact Hitters

Player-seasons: 2,190
Share of sample: 23.8 percent

The patient-contact group combines:

  • Above-average OBP
  • Above-average walk rates
  • Low strikeout rates
  • Approximately average power
  • Below-average stolen-base production

Its cluster center has an OBP score of:

z_{\mathrm{OBP}} = +0.55

and a strikeout-rate score of:

z_{\mathrm{SO/PA}} = -0.50

These hitters generally controlled the strike zone and put the ball in play. They were not necessarily powerless, but power was not the defining feature of the group.

The most statistically central examples include:

  • Justin Turner, 2022
  • Eric Hosmer, 2015
  • Edgardo Alfonzo, 1999

These are not necessarily the greatest seasons in the cluster. They are the seasons located closest to its statistical center.

2. Low-Impact Contact Hitters

Player-seasons: 1,925
Share of sample: 20.9 percent

This group also struck out infrequently, but without the OBP, walks, or power of the patient-contact cluster.

Its center was:

z_{\mathrm{OBP}} = -0.74 z_{\mathrm{ISO}} = -0.94 z_{\mathrm{BB/PA}} = -0.81 z_{\mathrm{SO/PA}} = -0.84

These hitters made contact, but much of that contact produced limited offensive value. Their low strikeout rates should not automatically be interpreted as evidence of superior hitting.

This distinction is important. Avoiding strikeouts is valuable only when the resulting balls in play produce enough hits, power, or advancement to compensate for the lost walks and extra-base production.

Central examples include:

  • Danny Bautista, 2004
  • Marlon Anderson, 2002
  • Melky Cabrera, 2007

3. Elite Power-and-Patience Hitters

Player-seasons: 851
Share of sample: 9.2 percent

This is the smallest cluster and the most offensively dominant.

Its center was approximately:

z_{\mathrm{OBP}} = +1.74 z_{\mathrm{ISO}} = +1.25 z_{\mathrm{BB/PA}} = +1.74

The cluster’s strikeout rate was almost exactly average relative to each season:

z_{\mathrm{SO/PA}} = +0.05

These hitters combined elite on-base ability with elite power and patience. Unlike the power-and-strikeout group, they did not require an exceptionally high strikeout rate to produce their power.

Central examples include:

  • Al Kaline, 1966
  • Kris Bryant, 2017
  • Ben Zobrist, 2009

The presence of players from widely separated eras is exactly what the season adjustment was designed to reveal. Their raw strikeout totals and league environments differed, but their relative offensive structures were similar.

4. Power-and-Strikeout Hitters

Player-seasons: 1,363
Share of sample: 14.8 percent

This group most closely resembles the familiar three-true-outcomes hitter.

Its defining characteristics were:

z_{\mathrm{ISO}} = +1.00 z_{\mathrm{BB/PA}} = +0.62 z_{\mathrm{SO/PA}} = +1.25

The group produced power and drew walks, but also struck out much more frequently than its seasonal peers.

Its OBP remained modestly above average:

z_{\mathrm{OBP}} = +0.25

Central examples include:

  • Jack Clark, 1979
  • Andruw Jones, 2002
  • Dale Murphy, 1986

This cluster existed long before the recent explosion in league-wide strikeouts. The modern environment made the raw statistical profile more common, but the relative archetype was already present.

5. Aggressive Free-Swingers

Player-seasons: 1,946
Share of sample: 21.1 percent

The aggressive free-swinging group had:

  • Below-average OBP
  • Below-average walk rates
  • Above-average strikeout rates
  • Approximately average power
  • Little baserunning contribution

Its center included:

z_{\mathrm{OBP}} = -0.83 z_{\mathrm{BB/PA}} = -0.66 z_{\mathrm{SO/PA}} = +0.66

Power was only slightly above the seasonal average:

z_{\mathrm{ISO}} = +0.06

This is an important contrast with the power-and-strikeout cluster. Both groups struck out frequently, but the aggressive free-swingers did not receive the same compensating power, walks, or OBP.

Central examples include:

  • Ollie Brown, 1969
  • Ryan Ludwick, 2009
  • Matt Williams, 1998

6. Speed-and-Contact Hitters

Player-seasons: 943
Share of sample: 10.2 percent

This was the most specialized cluster.

Its net stolen-base score was:

z_{\mathrm{NetSB/PA}} = +2.20

No other cluster approached that level.

The group also had:

z_{\mathrm{SO/PA}} = -0.32

and:

z_{\mathrm{ISO}} = -0.69

These players produced value through speed, contact, and mobility rather than power.

Central examples include:

  • Delino DeShields, 2000
  • Stan Javier, 1995
  • Marquis Grissom, 1994

Their average OBP was almost exactly equal to the seasonal mean. The cluster was not defined by superior hitting in the narrow sense. A distinctive combination of speed, contact, and limited power defined it.

Did the Archetypes Change Over Time?

This was the study’s central question.

The raw offensive environment changed dramatically. Figure 2 shows a major increase in strikeout rates, higher power levels, and shifting patterns of stolen bases.

The relative distribution of player archetypes was surprisingly stable.

Figure 6. Five-year moving percentages of qualified player-seasons assigned to each cluster.

Patient-contact hitters generally represented between approximately 20 and 28 percent of qualified seasons.

Low-impact contact hitters reached their highest levels during the 1960s and 1970s, then gradually declined. Their five-year moving share peaked near 26 percent in the mid-1970s and stood near 19 percent by 2025.

Power-and-strikeout hitters became somewhat more prevalent, reaching approximately 18 percent around 2012. Yet they never displaced the other archetypes.

Speed-and-contact hitters remained surprisingly consistent. Their moving share generally remained between approximately 8 and 12 percent, even though league-wide stolen-base environments changed considerably.

The elite power-and-patience group remained rare throughout the entire study. It accounted for approximately 8 to 11 percent of qualified seasons in most periods.

The 2020s did not produce a completely new distribution of player types. Compared with the partial 1950s sample, the modern distribution contains somewhat fewer low-impact contact hitters and a modestly larger share of power-and-strikeout hitters.

The basic architecture, however, remains recognizable.

This suggests that baseball evolution has operated on at least two levels.

At the first level, the statistical baseline changes. Strikeouts become more common. Power becomes more valuable. Stolen-base strategies change.

At the second level, players continue to occupy recurring roles relative to that baseline. Every era still contains:

  • Patient hitters
  • Free swingers
  • Power hitters
  • Speed specialists
  • Low-impact contact hitters
  • Rare players who combine several elite skills

The numbers change. The ecological niches persist.

Offensive Archetypes and Defensive Positions

The clusters were also closely connected to defensive position.

Figure 7. Primary defensive positions represented within each offensive archetype.

The elite power-and-patience cluster was concentrated at traditional offensive positions:

  • 28.2 percent first basemen
  • 17.7 percent right fielders
  • 14.1 percent left fielders
  • 12.5 percent third basemen

Shortstops represented only 2.1 percent of the cluster.

The power-and-strikeout group followed a similar pattern. First base, right field, and left field accounted for a large portion of those seasons.

The speed-and-contact cluster looked completely different:

  • 32.1 percent center fielders
  • 20.0 percent shortstops
  • 19.7 percent second basemen
  • 15.0 percent left fielders

Catchers and first basemen were almost absent.

The low-impact contact group was strongly concentrated in the middle infield:

  • 27.9 percent shortstops
  • 23.9 percent second basemen

This pattern reflects the interaction between offense and defensive value. A shortstop could remain in a lineup with limited power because his defensive position carried different offensive expectations. A first baseman generally needed much greater offensive production.

The archetypes are therefore not purely hitting categories. They also reflect the way teams distribute offensive and defensive responsibilities across the field.

What the PCA Map Really Shows

The PCA figure is not simply a picture of six boxes.

Instead, it shows a continuous offensive landscape.

The elite power-and-patience hitters occupy the high end of the first principal component because they combine OBP, power, and walks.

The power-and-strikeout hitters move upward on the second component because of their combination of ISO and strikeout rate.

The speed-and-contact hitters move in the opposite direction because their offensive identities are dominated by baserunning and lower power.

The patient-contact and low-impact contact groups overlap along the contact dimension, but separate sharply through OBP and walk rate.

This provides a useful reminder about classification. A player’s archetype is not his complete identity. It is a summary of how his season relates to thousands of other seasons in baseball history.

Some players sit close to a cluster center. Others occupy transitional areas between types.

Limitations

This study has several important limitations.

First, the model uses traditional Lahman statistics. It does not include park-adjusted measures such as wRC+, nor does it include Statcast measures such as exit velocity, barrel rate, launch angle, or sprint speed.

Second, the model treats each player-season as an independent observation. A player who qualified in 15 seasons appears 15 times. This is appropriate for studying the distribution of seasonal styles, but it gives durable players more influence than short-career players.

Third, the cluster labels are interpretations. The algorithm identifies groups of seasons that are statistically similar. It does not name those groups.

Fourth, the sample includes only qualified hitters. Part-time players, platoon specialists, defensive replacements, and many late-career seasons are excluded.

Fifth, seasonal standardization intentionally removes changes in the league-wide baseline. This is the correct approach for identifying relative archetypes, but it means the clusters should not be interpreted as absolute comparisons of offensive production.

A power-and-strikeout hitter from 1960 did not necessarily strike out as frequently as a member of the same cluster in 2025. He struck out frequently relative to the hitters around him.

Conclusion

I began this study expecting to find a succession of offensive types.

I expected contact hitters to dominate the early years, speed players to expand during the 1970s and 1980s, and power-and-strikeout hitters to overwhelm the modern period.

Some of those movements are visible, but the broader result is more interesting.

Baseball’s offensive environment changed enormously. Its fundamental player archetypes changed much less.

The low-strikeout hitter did not disappear. His raw strikeout rate simply rose with the league.

The power-and-strikeout hitter did not suddenly appear in the twenty-first century. Earlier versions existed within lower-strikeout environments.

The speed-and-contact player did not vanish during the power era. His share remained more stable than the raw stolen-base totals might suggest.

Perhaps the evolution of the baseball player is not a story of one species replacing another.

Perhaps it is a story of persistent roles adapting to a changing environment.

The game changes its equilibrium. The players reorganize themselves around it. Yet the same broad offensive strategies continue to reappear, season after season and generation after generation.

That continuity may be one of the most striking features of baseball history.