Reliability, Redundancy, and Standards
A system that needs heroes to keep working is a system that has not been engineered.
A system that needs heroes to keep working is a system that has not been engineered.
Every company that has run for a while has a person like this, and one company I think of had a particularly clear example. The growth held together because of one operator who seemed to be everywhere at once, who knew which numbers to watch and when they were lying, who could feel a channel starting to slip before any report showed it and intervene in time, who carried in their head the whole working knowledge of how the system actually ran. The results were excellent, and everyone admired the operator, and they were right to, because the work was truly impressive. Then the operator burned out and left, and within two quarters the growth that had looked like a machine revealed itself to have been a person, and the company spent a year and a great deal of money discovering that it had never had a system at all, only someone heroic enough to act like one.
The lesson the company drew was the wrong one, which was to go and find another hero. They hired for brilliance, looked for someone with the same rare instincts, and set about recreating their dependence on an exceptional individual, because the idea that the problem was the dependence itself, rather than the loss of the particular person, never quite surfaced. They had confused a staffing problem with a structural one, and the structural one would recur with every hero they found, because a system built to need a hero is always one hero away from collapse.
This is a particular kind of failure, and it points at a particular kind of absence. The company had results without reliability, outcomes that depended entirely on an exceptional individual and disappeared when the individual did, and the difference between that arrangement and an engineered one is the subject worth examining, because it is the difference between a practice that gets lucky with its people and a discipline that produces reliable outcomes from ordinary ones. Reliability, the property of continuing to work under normal variation and without heroics, is something engineering learned to manufacture deliberately, through a small set of methods that go-to-market has almost entirely failed to adopt, and the result is a field that remains dependent on heroes precisely because it never built the structures that would let it stop being.
Reliability is manufactured, not hoped
The first thing to understand is that reliability in engineering is a designed property, something built into a system on purpose through known techniques, with its own mathematics and its own methods. An engineer designs a bridge to a specified reliability, with the failure modes understood and guarded against, so that staying up is a property the structure has by construction rather than a fortunate outcome it happens to produce or a thing anyone is left hoping for. Reliability engineering, as a field, exists precisely to make this systematic, to turn the continued working of a system from a hope into a specification.
It is worth seeing that reliability is a property of the whole system over time, and that it can be described precisely. You can write the reliability of a system as the probability that it is still working at a given time, a number that starts near one and declines as the chance of failure accumulates, and the shape of that decline, how fast reliability falls and why, is exactly what reliability engineering studies and designs against. The discipline has a detailed account of how systems fail over their lifetimes, the early failures from defects, the steady middle period, the wear-out failures at the end, and it uses that account to design systems that achieve a stated reliability over a stated life. The point is that reliability is a quantity you can specify, design for, and verify, rather than a vague aspiration that some systems happen to satisfy.
The shape of failure over time even has a familiar form, sometimes called the bathtub curve, high failure early as initial defects surface, a long low-failure middle once the survivors have proven themselves, and rising failure at the end as things wear out. Engineers use the shape to decide where to invest, burning in components to clear the early failures, planning replacement before the wear-out period begins. Go-to-market systems have their own version of this lifecycle, the shakeout of a new motion, the stable middle of a working one, the slow decay as a channel ages and an audience tires, and a field that thought in these terms would anticipate each phase rather than being surprised by all three. The verb that matters underneath all of it is verify. A reliability claim in engineering is something you can test, by running components to failure, by modeling the system, by checking the design against the code, so that the stated reliability is backed by evidence rather than asserted. Go-to-market makes no reliability claims at all, and so has nothing to verify, and the absence stays invisible precisely because the field has never thought to ask the question that would expose it, which is simply how reliable is this system, stated as a number and defended with evidence.
Go-to-market has almost no equivalent of any of this. It does not specify the reliability of its systems, does not design for a stated probability of continuing to work, does not study how its systems fail over time in any systematic way, and as a result it produces reliability, when it produces it at all, by accident or by heroics. The company with the heroic operator had no reliability in the engineering sense, because the continued working of its growth was a personal achievement of one individual rather than a designed property of the system, and a personal achievement does not amount to a reliability, because it leaves with the person. A field that wanted reliable outcomes would have to learn to manufacture reliability the way engineering does, and the methods for doing so are well known and almost entirely unused here.
Redundancy
The first of those methods is redundancy, the deliberate provision of more than one way for the system to do its essential work, so that the failure of any single part does not bring the whole system down. Redundancy is the direct answer to the single point of failure, the dependence on one component whose loss is catastrophic, and it is one of the most powerful techniques engineering has, because it can produce a system far more reliable than any of its parts.
The mathematics of why is simple and worth seeing. Suppose a system depends on a single channel that works eighty percent of the time, so its reliability is four fifths and it fails one time in five. Now suppose instead the system has two independent channels, each working eighty percent of the time, arranged so that either one alone is enough. The system fails only if both fail, and if the channels are truly independent, the chance of both failing is one fifth times one fifth, one twenty-fifth, so the system now works twenty-four times out of twenty-five, a reliability of ninety-six percent built out of two components that were each only eighty percent reliable. Redundancy turns unreliable parts into a reliable whole, and the improvement is dramatic, which is why every system that has to be reliable, from aircraft to data centers, is built with it.
The crucial word in all of that is independent, and it is where the technique is most often botched. Two channels that fail for the same underlying reason are not redundant, however separate they look, because the thing that takes one down takes the other with it. A company running two paid channels on the same platform has less redundancy than it imagines, since a change in the platform’s rules can fail both at once, and the apparent backup vanishes exactly when it is needed. Real redundancy requires the parts to fail for different reasons, which is harder and more expensive than simply having more than one of something, and the discipline is in securing true independence rather than the comforting appearance of it. A reliable demand system has paths that draw on different audiences, different mechanisms, and different dependencies, so that no single event can take them all down together.
A worked version in the field’s own terms shows the idea concretely. Suppose a company has two ways of generating qualified demand, an inbound content engine and an outbound motion, and suppose each one alone delivers its target in a given quarter about eighty percent of the time, failing one quarter in five to some cause of its own, a content-ranking shift for the one, a deliverability problem for the other. Run only one of them and the company misses its number one quarter in five. Run both, and as long as their failure causes are truly independent, the company misses only when both happen to fail in the same quarter, which occurs about one time in twenty-five, so the reliability of hitting the number rises from eighty percent to ninety-six percent. The reliability of the whole, written as the probability that the system still delivers at a given time, is built up out of parts that are each far less reliable, and the gain came entirely from the structure rather than from making either path better. The condition that makes it work is the independence, the ranking shift and the deliverability problem having nothing to do with each other, and a company that built both its inbound and its outbound on the same platform would forfeit most of the gain, because the single event could take both down at once.
Go-to-market, left to its own instincts, builds the opposite, because redundancy looks like waste to a field trained to optimize. The locally efficient move is always to concentrate, to pour everything into the one channel that works best, the one message that converts, the one motion that is winning, and concentration is the direct production of single points of failure, a system with no redundancy and therefore no reliability, excellent until the one thing it depends on fails. The heroic operator was a single point of failure in human form, the whole system’s essential knowledge and judgment concentrated in one person with no redundancy, and the company optimized for the efficiency of having one brilliant person carry everything until the day that efficiency revealed its cost. A discipline would build redundancy on purpose, accepting the apparent inefficiency as the price of reliability, because it would understand that a system with no redundancy is a system waiting for its single point of failure to find it.
Standards and codes
The second method is the one that does the most to make a field reliable, and it is the one go-to-market most conspicuously lacks, which is standards. A mature engineering discipline carries a body of codified standards, building codes, design specifications, accepted methods, written down and shared across the whole field and updated as knowledge accumulates, and these standards are most of what lets ordinary practitioners produce reliable results. An engineer designing a structure works within codes that encode the accumulated knowledge of the field, including the lessons of past failures, rather than deriving everything from first principles and hoping, so that the floor of competent practice is high and the typical building is safe even when no genius was involved.
This is the deepest difference between a discipline and a practice, and it is worth dwelling on. In a discipline with standards, the knowledge lives in the field rather than in the individuals, written into codes and methods that any competent practitioner can apply, so that the field as a whole gets reliably good results without depending on the rare brilliance of particular people. The standards are how a discipline makes its hard-won knowledge portable and durable, how it ensures that a lesson learned once is applied everywhere and not relearned at full cost by each new practitioner, and how it raises the floor so that ordinary competence produces good outcomes. A field with strong standards does not need heroes, because the standards do the work that heroes would otherwise have to do.
It helps to be concrete about what a standard does. A standard raises the floor rather than the ceiling, it makes the typical result reliably good without making the best result better, and that is precisely its value, because a field is mostly made of typical practitioners and typical work, and lifting that vast middle matters more for the field’s overall reliability than enabling a few peaks. A building code produces buildings that do not fall down rather than beautiful architecture. A go-to-market standard, if the field built one, would make the ordinary effort reliably competent rather than making a brilliant campaign more brilliant, encoding the known ways to size a channel, to build in margin, to test a tolerance before scaling, to avoid the documented failure modes, so that a team of ordinary people following the standard would steer clear of the disasters that ordinary teams currently walk into for lack of any shared guidance.
The difference between a playbook and a standard is worth making sharp, because the field has the former and mistakes it for the latter. A playbook is a description of what worked for someone, shared in the hope that it works again, carrying no guarantee and tied to the conditions of its origin, which is why playbooks expire when the conditions change. A standard is a codified method the field has tested and established, maintained and updated as knowledge grows, carrying the weight of collective verification rather than individual anecdote. The playbook says here is what we did. The standard says here is what works, here is the evidence, and here is how to apply it, and the gap between those two kinds of statement is the gap between a practice trading tips and a discipline accumulating knowledge.
Go-to-market has nothing of the kind. It has playbooks, which churn with the tools and encode no durable knowledge, and it has fashions, which spread and fade, and it has the private methods of individual practitioners, which leave when they do, and it has no shared, codified, accumulating body of standards that would let ordinary practitioners produce reliable results. Each team reinvents its methods from the material at hand, each new practitioner relearns the lessons at full cost, and the knowledge that should accumulate in the field instead evaporates with the people who held it, which is why the field stays permanently dependent on finding rare talent. The absence of standards is the absence of a mechanism for the field to know what it knows, and without that mechanism reliability is impossible, because reliability is exactly the property of knowing that a method will work because the field has established that it does.
The cult of the hero
There is a cultural symptom of all this worth naming, because the field mistakes it for a virtue, which is the celebration of heroics. Go-to-market loves its heroes, the operator who saved the quarter, the founder who pulled growth out of nothing through sheer will, the marketer whose instincts no one can replicate, and it tells these stories with admiration and holds them up as the ideal. The admiration is understandable, because the work is real and impressive, and the heroics truly do save quarters. The trouble is that a system that requires heroics to keep working is, by that very fact, an unreliable system, and the prominence of heroes is a symptom of the absence of engineering rather than a sign of its presence.
Mature engineering disciplines have a different relationship to heroism, one worth importing. They honor the heroic recovery, the pilot who lands the failing aircraft, the engineer who catches the flaw at the last moment, and at the same time they understand that every such episode represents a failure of the system to be reliable without heroism, and they work relentlessly to engineer away the need for it. The goal of aviation safety is aircraft and procedures so reliable that bravery is rarely required, rather than a supply of braver pilots, and the field measures its progress by how seldom heroism is needed rather than by how often it is displayed. Go-to-market has the relationship backwards, treating the need for heroes as a permanent feature to be staffed for rather than a deficiency to be engineered out, and as long as it does, it will keep building systems that work brilliantly until the hero leaves and then do not work at all.
There is also a human cost to the arrangement that the romance obscures. A system that runs on heroics runs on the heroes, and heroes burn out, because the role of holding an unreliable system together by personal effort is not survivable indefinitely, and the field’s dependence on heroism is therefore also a quiet machine for using people up. The operator who held the company together left for a reason, because the work of being a system was exhausting in a way no amount of admiration compensated for, and the company that loses its hero to burnout has usually spent the person rather than been abandoned by them. A field that engineered reliability would be kinder to its people almost as a side effect, because it would stop asking individuals to be the redundancy and the standard and the reliability all at once.
What the infrastructure would require
Reliability through standards is not something a single team can fully build for itself, because standards are by their nature shared, which raises the harder question of what it would take for go-to-market to develop the institutional infrastructure a discipline requires. Engineering reliability rests on shared institutions, the bodies that write and maintain the codes, the accumulated literature of failures and methods, the training that transmits the standards, the professional norms that hold practitioners to them, and go-to-market has almost none of this shared infrastructure, which is part of why its knowledge will not accumulate.
Building it would be a long project, the work of turning a practice into a profession, and it would require things the field has never had, a shared account of the object being engineered so that standards have something to be standards about, a culture of studying and recording failures so that the codes have something to encode, and institutions willing to maintain the shared knowledge across the churn of companies and tools. None of this is impossible, and other fields have done exactly it, moving from collections of talented individuals to professions with shared standards over the course of decades, but it requires the field to want reliability more than it wants the romance of heroics, which is a real cultural choice and not a foregone one.
The histories are worth knowing, because they show the move is possible. Medicine was, over the last century or so, much closer to a craft of gifted individuals than the standardized profession it has become, and it crossed the distance through specific institutional acts, the standardization of training, the licensing that set a floor, the accumulation of a shared literature, the professional bodies that maintained the standards and held practitioners to them. Engineering itself professionalized the same way, turning the master builder’s private knowledge into shared codes maintained by institutions, with credentials that certified a practitioner had absorbed the standard. None of this happened on its own, and none of it was free, and in each case the field had to decide that reliable competence from many was worth more than occasional brilliance from a few, which is exactly the decision go-to-market has not yet made. The reward would be a field that produces good outcomes reliably from ordinary practitioners, which is what a discipline is, and the cost would be giving up the story the field most loves to tell about itself.
There is a question underneath the question of reliability, and it grows more insistent as the systems get more powerful. A field that learns to build reliable, redundant, standardized systems for creating and capturing demand will have built something truly powerful, machines that work dependably at scale on the attention and behavior of large numbers of people. A system can be made perfectly reliable at doing something no one should be doing reliably, and the better the field gets at the first problem, the harder it becomes to keep avoiding the second.
References
- On reliability engineering, the reliability function, hazard rates, and the design of redundant systems, see standard treatments in reliability and systems engineering.
- On standards, codes, and the role of codified knowledge in raising the floor of practice, see the history of engineering standardization and professional codes.
- On the progression from a collection of skilled individuals to a profession with shared standards, the histories of structural engineering and of medicine offer the clearest parallels.