Industrial reliability lessons from Pearl Harbor and Midway
Key Highlights
- Reliability involves managing uncertainty, not eliminating it.
- Historical examples like Pearl Harbor and Midway demonstrate the power of converting uncertainty into actionable information.
- Effective reliability practices include continuous observation, interpretation, decision-making, and adaptation to changing conditions.
- Complex systems depend on critical dependencies that, if managed properly, can prevent failures before they occur.
Management of Uncertainty is a new Plant Services column from contributor Michael D. Holloway. This new column focuses on a simple premise: reliability is not the elimination of uncertainty, but rather disciplined management. Every maintenance strategy and practice is ultimately an attempt to understand what we know, recognize what we do not know, and make better decisions before uncertainty becomes a consequence. The practical issues explored in each article will examine how information, evidence, risk and judgment can reduce uncertainty and improve reliability.
Reliability is often described as the probability that a machine, component or system will perform its intended function for a specified period of time under stated conditions. While that definition is mathematically useful, I have increasingly come to believe that it describes the measurement of reliability more effectively than it describes the discipline itself. Reliability, in practice, is something broader.
Pearl Harbor exposed the power of managed uncertainty
Few historical events demonstrate that principle more dramatically than the six months separating Pearl Harbor from the Battle of Midway. During that remarkably short period the balance of capability in the Pacific began to change, not simply because one side acquired better ships, aircraft or weapons, but because one organization became increasingly effective at converting uncertainty into information, information into decisions and decisions into action.
On December 7, 1941, Japan demonstrated what happens when uncertainty is successfully managed and simultaneously imposed upon an opponent. The Japanese knew where its carrier force was, knew what it intended to do, understood its timetable, protected the information surrounding the operation, and constructed an enormous logistical and operational system capable of moving six aircraft carriers and their supporting vessels thousands of miles across the Pacific without revealing their purpose. The Americans, meanwhile, possessed fragments of information but lacked the integrated understanding necessary to convert those fragments into effective action. The result was catastrophic for the United States.
Operation K revealed critical information
Only 87 days later, during the night of March 3–4, 1942, Japanese aircraft returned to Hawaii in an operation that has largely disappeared from popular memory. Known as Operation K, the mission employed two enormous Kawanishi H8K flying boats that traveled from the Marshall Islands to French Frigate Shoals, where they rendezvoused with Japanese submarines for refueling before continuing toward Oahu. The mission was extraordinary from a logistical perspective, but it was also extraordinarily dependent upon assumptions.
Aircraft performance had to be sufficient. Their navigation had to be accurate. Of course the weather had to cooperate. The submarines had to reach the correct location and refueling had to succeed. During all of this, the communications had to remain secure. The Americans had to remain unaware of the refueling arrangement, and once the aircraft reached Hawaii, crews had to locate their targets in darkness. Several of those assumptions failed.
Cloud cover over Oahu prevented accurate observation and targeting, and bombs fell harmlessly or caused only minor damage. Perhaps most importantly, the mission provided the Americans with something considerably more valuable than the physical damage Japan hoped to inflict; it provided information.
Information becomes valuable when it changes decisions
American intelligence did not simply observe that Japanese aircraft had returned to Hawaii and conclude that another attack had occurred. Analysts began asking the question that every reliability professional should ask when confronted with an unexpected event: What system made this possible?
That question changes everything because once the event is treated as the output of a system rather than as an isolated occurrence, the investigation moves upstream. Where did the aircraft originate? How could they possess sufficient range? Where were they refueled? What communications preceded the operation? What supporting assets were required? Which dependencies had to remain intact for the mission to succeed?
Data from the event, when correctly interpreted, reduces uncertainty.
By May 1942, Japan was preparing an operation against Midway Island designed in large part to draw the remaining American aircraft carriers into a battle in which they could be destroyed. For the Japanese plan to work as intended, however, commanders needed information about American carrier movements, particularly whether those carriers remained at Pearl Harbor. Japan, therefore, planned another long-range reconnaissance mission using essentially the same refueling arrangement at French Frigate Shoals. But this time around the Americans understood the system.
American intelligence had been reading enough Japanese communications to recognize the importance of the location, and Admiral Chester Nimitz ordered American vessels to patrol the area. When Japanese submarines arrived to support the planned reconnaissance operation, they discovered American ships occupying the refueling point, making the mission impractical and forcing its cancellation.
No great naval engagement occurred at French Frigate Shoals. No aircraft carriers exchanged fire and no battleships were sunk; however, something extraordinarily important happened, a critical dependency had been identified and interrupted. This is where Operation K stops being merely an interesting footnote to World War II and becomes a lesson in reliability engineering.
Every complex system contains dependencies, and some of those dependencies possess importance vastly disproportionate to their physical size or apparent significance. A manufacturing facility may contain thousands of components, hundreds of motors, dozens of pumps, sophisticated automation, highly trained personnel, and millions of dollars of equipment, yet the capability of that entire system may ultimately depend upon a small bearing receiving the correct lubricant, a seal maintaining integrity, a cooling circuit remaining unobstructed, a sensor providing trustworthy information, or an operator recognizing a developing condition before it crosses the threshold into failure. The reliability professional therefore should not merely ask what can break. The better question is: Where does uncertainty enter the system, upon what assumptions does capability depend, and what information would allow us to intervene before uncertainty becomes failure?
That is precisely what happened between Pearl Harbor and Midway. The Japanese system that achieved extraordinary operational surprise in December 1941 had not suddenly become incompetent six months later, the Americans had learned the system and therefore, could influence it.
Observation produced information. Information was interpreted. Interpretation informed decisions. Decisions produced actions. Outcomes were measured.
Those outcomes created additional information, and the cycle began again.
Reliability is a cycle of learning and adaptation
Uncertainty → Observation → Information → Interpretation → Decision → Action → Measurement → Learning → Adaptation → Capability
This progression is as applicable to an industrial plant as it is to military operations, because capability does not emerge simply from possessing equipment, technology, or people. Capability emerges when a system can observe its environment, distinguish meaningful signals from noise, understand what those signals imply, make appropriate decisions, measure the consequences, and modify future behavior accordingly.
This is also why condition monitoring, inspection, lubrication analysis, vibration analysis, process data, maintenance history and operator observations have value far beyond the individual measurements they produce. Their purpose is to reduce uncertainty sufficiently that someone can make a better decision before capability is lost. The absence of information does not eliminate uncertainty, it merely forces the organization to substitute assumptions for knowledge.
Japan approached Midway carrying several critical assumptions about where the American carriers were and how the Americans would respond. The failed reconnaissance operation meant that some of those assumptions could not be independently verified. The United States, meanwhile, had developed enough intelligence about Japanese intentions to position its carriers northeast of Midway before the Japanese strike force arrived. The result is well known.
On June 4, 1942, American carrier aircraft struck the Japanese fleet, and by the end of the battle Japan had lost four fleet carriers that had participated in the attack on Pearl Harbor. The strategic balance in the Pacific had begun to turn.
It is tempting to describe Midway primarily as a story of courage, tactics and fortunate timing, and all of those elements unquestionably mattered, but underneath them was another contest that reliability professionals should immediately recognize. It was a contest between two systems attempting to manage uncertainty.
One system entered the battle increasingly dependent upon assumptions it could not verify. After earlier defeat, the other had built an increasingly effective mechanism for collecting information, interpreting weak signals, challenging assumptions, and positioning capability before the moment of demand arrived. That is what reliability ultimately seeks to accomplish.
Reliability means staying ahead of uncertainty
We cannot eliminate uncertainty from machines, manufacturing plants, electrical grids, transportation networks, supply chains, organizations, or military operations because uncertainty is inherent in every complex system. Components degrade, environments change, measurements contain error, people make mistakes, assumptions become obsolete and conditions emerge that designers never anticipated. The objective is therefore not to eliminate uncertainty. The objective is to manage it better than the system can surprise us.
Pearl Harbor demonstrated the extraordinary capability that can be exposed when uncertainty is successfully managed and imposed upon an opponent. Midway demonstrated what can happen when the informational advantage reverses, and one organization begins learning faster than another.
Reliability is the disciplined management of uncertainty to expose capability.
And sometimes the most consequential reliability action is not repairing what failed, it is understanding the system well enough to prevent uncertainty from deciding what happens next.
About the Author
Michael D. Holloway
5th Order Industry
Michael D. Holloway is President of 5th Order Industry which provides training, failure analysis, and designed experiments. He has 40 years' experience in industry starting with research and product development for Olin Chemical and WR Grace, Rohm & Haas, GE Plastics, and reliability engineering and analysis for NCH, ALS, and SGS. He is a subject matter expert in Tribology, oil and failure analysis, reliability engineering, and designed experiments for science and engineering. He holds 16 professional certifications, a patent, a MS Polymer Engineering, BS Chemistry, BA Philosophy, authored 12 books, contributed to several others, cited in over 1000 manuscripts and several hundred master’s theses and doctoral dissertations.
