Automotive service advisor presenting a repair estimate while a performance dashboard displays sales targets and recommendation metrics.
Article

When a Metric Becomes an Instruction: What Sears Auto Centers Reveals About Performance Measurement

When I first examined the Sears Auto Centers case, the sales figures were not the most interesting part. What mattered was what happened after those figures became targets.

Sears was not only measuring how much employees sold. The performance measurement system surrounding those sales—commissions, performance comparisons, contests, and product-specific goals—also communicated what the organization expected employees to prioritize. That distinction matters because a number can be accurate while still producing a misleading picture of performance.

Repair revenue can be recorded correctly. Sales per hour can be calculated correctly. Product targets can be tracked correctly. None of those measurements, however, can determine whether every recommended repair was necessary, appropriate, or aligned with the customer’s interest.

The central problem was therefore not whether Sears collected accurate data. It was whether the data represented the outcome the company actually needed to protect.

A higher repair sale could indicate that an employee had correctly identified a legitimate mechanical problem. It could also indicate that more work had been recommended than the customer required. The recorded number would look the same in both situations because the sale was measurable while the necessity behind it was harder to observe.

Once compensation and performance pressure were attached to the measurable part, the metric stopped functioning only as a description of activity. It began influencing the activity itself.

This is the principle I find most important in the case: a performance metric becomes an instruction when rewards, recognition, or employment depend on it. When that metric is only an incomplete proxy for the real objective, employees can improve the number while weakening the outcome the organization intended to protect.

A Legitimate Signal Became an Incomplete Definition

It would be easy to look back at Sears’ compensation system and treat the entire idea as obviously irrational. That would make the case simpler, but it would also make the analysis weaker.

Sales were a legitimate business signal. Sears operated automotive service centers where employees inspected vehicles, interpreted mechanical conditions, recommended repairs, and sold parts and services. Revenue could reasonably help management evaluate productivity, demand, employee activity, and store performance.

Higher sales could indicate that advisors were identifying real problems, explaining them clearly, and helping customers authorize necessary work. A store generating more repair revenue might be serving more customers, completing more jobs, or operating more efficiently. There was nothing inherently meaningless about the number.

The difficulty emerged when sales began carrying more interpretive weight than the metric could support.

A repair sale does not explain why the transaction occurred. It does not show whether the recommendation was urgent, preventive, optional, unnecessary, or misunderstood. It does not reveal how confident the customer was in the decision, nor whether the advisor protected the long-term relationship or simply increased the immediate transaction. The metric contained useful information, but it did not contain the entire truth of the service interaction.

This is where I think organizations often make a critical mistake. A metric begins as one signal among several, then gradually becomes the practical definition of performance.

At Sears, repair sales were reinforced through a wider system of incentives and expectations. Contemporary reporting described commissions, weekly sales tracking, employee comparisons, contests, rewards, goals connected to specific products and services, and an expectation of approximately $147 in sales per hour. Individually, each mechanism might have appeared manageable. Together, they created a consistent message: sell more, sell more frequently, prioritize the products the system had chosen to track, keep pace with the store average, and avoid falling behind employees whose numbers were higher.

Management may have intended these mechanisms to encourage productivity, but employees experienced them as operating conditions. That difference between managerial intention and employee interpretation is central to the case.

What the System Actually Taught Employees

Organizations often speak about metrics as though they remain passive after they are introduced — a dashboard records performance, a target clarifies expectations, a commission rewards success, a contest creates motivation. In practice, these mechanisms do more than observe behaviour. They define which behaviour is visible, valued, and rewarded.

Once employees understand that a number affects earnings, status, recognition, or job security, they begin adapting their decisions around it. That adaptation does not necessarily require deliberate dishonesty. An employee may simply become more attentive to opportunities that increase the measured result — recommendations growing more assertive, optional work framed as more important, ambiguous situations resolved in favor of action rather than restraint. Products connected to specific goals may receive more attention than products that are not being tracked.

The employee does not need to consciously decide to manipulate the system. The incentive can gradually influence judgment.

This is why I do not think a KPI should be evaluated only by asking whether it measures something real. Sales were real, revenue was real, and the repair orders were real. The more important question is what the organization taught employees to do in order to improve those numbers.

A metric can begin as a representation of behaviour and then become a cause of behaviour. Once that happens, the organization is no longer observing an independent signal but a result partly produced by the measurement system itself. That is where interpretation becomes difficult.

Management may see rising sales and conclude that productivity has improved. The increase may also reflect stronger selling pressure, more aggressive recommendations, or a shift in judgment caused by the incentive. The number alone cannot separate those explanations. The metric remains accurate. The interpretation becomes incomplete.

What I find significant is that the distortion does not need to appear inside the calculation. The arithmetic can remain correct while the surrounding behaviour changes. The organization may continue receiving clean reports, consistent totals, and apparently comparable performance data, yet the meaning of those numbers has shifted because employees are now responding to the measurement system itself.

The metric is no longer only recording the environment. It is helping create it.

The Risk Inside an Advisory Relationship

The Sears case was especially vulnerable to this problem because automotive repair is not a standard retail transaction.

A customer buying a visible product can often compare its price, quality, and features. The customer may not possess perfect information, but the object being purchased can usually be inspected and understood. Automotive service depends much more heavily on expert interpretation. Most customers cannot independently determine whether a brake component should be replaced, whether an alignment is necessary, whether a shock absorber is worn, or whether a repair can safely be delayed. They depend on the advisor to translate technical conditions into an understandable recommendation.

This creates a significant information imbalance.

The employee is functioning not only as a seller but as an interpreter of the customer’s need. I think this distinction changes how the incentive system should be evaluated, because the advisor is not merely encouraging the customer to purchase more of a product the customer already understands — the advisor is helping define what the customer believes is necessary. That role carries a different level of responsibility because the customer cannot easily verify the recommendation before accepting or rejecting it. When the same employee can benefit financially from recommending more work, the system creates a structural conflict.

That does not prove that every employee acted improperly. It does not establish that every recommendation was unnecessary, and it does not mean commissions always produce misconduct. It means the incentive was placed inside a relationship where the customer had limited ability to evaluate the advice independently.

The risk was therefore not only that an employee might choose to oversell — the deeper risk was that the system could shift the threshold for what employees considered worth recommending. A repair that once appeared optional might begin to look advisable. A preventive service might be presented with greater urgency. A borderline condition might be resolved in favor of replacement. The distinction between service and selling becomes harder to maintain when a system rewards one side of the decision more clearly than the other.

Sales were visible; restraint was not.

A recommendation generated revenue and appeared in the performance data. A decision not to recommend unnecessary work protected the customer, but it produced no equivalent metric. That imbalance matters because organizations tend to reward what they can observe. When responsible restraint is difficult to measure, it can disappear from the formal definition of performance even when it remains essential to the real outcome.

From a data-interpretation perspective, this is one of the most important features of the case. The metric did not merely omit part of the service interaction. It omitted the part that could not easily be converted into a transaction. The system could see what was sold, but it could not see the value of a recommendation that was never made.

Allegations, Denials, and the Limits of the Evidence

The Sears Auto Centers case attracted regulatory scrutiny in 1992. California authorities alleged that Sears had charged customers for unnecessary repairs and argued that the compensation system encouraged employees to recommend work customers did not need. Similar concerns and investigations emerged in other states.

Sears disputed that characterization. The company denied that it had encouraged systematic overselling and challenged the methods used by investigators. It argued that some of the targets described as quotas were intended to reflect workload expectations and that individual incidents did not prove a company-wide pattern of fraud.

Those distinctions need to remain visible.

I do not think the case should be reduced to a simple claim that Sears established sales targets and every employee responded by selling unnecessary repairs. The evidence does not support such a universal conclusion. The regulatory allegations, the company’s denials, and the limits of what can be proven about individual behaviour should remain separate.

But the measurement problem does not depend on proving that every allegation was correct. The system can be analyzed on its own terms. Sears tied compensation and performance pressure to repair sales within a service relationship defined by information asymmetry. The metric could record the value of the transaction, but it could not establish whether the recommendation behind that transaction was necessary. That alone created a vulnerability.

The strongest evidence came not only from regulators or critics, but from Sears’ own leadership.

Sears chairman Edward Brennan continued to deny that the company had engaged in systemic fraud. At the same time, he acknowledged that incentive compensation and goal setting had created an environment in which mistakes could occur. He pointed to aggressive selling and rigid attention to goals as possible explanations and accepted management responsibility for the policies.

I find that acknowledgment decisive because it moves the case beyond a disagreement over isolated events. Management recognized that the operating environment could influence employee judgment. The problem was not framed only as a matter of individual misconduct. It was also treated as a consequence of the system surrounding the employee. That is a more useful interpretation because it directs attention toward the structure that shaped behaviour.

When an organization responds to a performance failure only by blaming employees, it avoids examining the instructions embedded in its own metrics. Employees still remain responsible for their decisions, but management is responsible for designing the conditions under which those decisions are made, measured, compared, and rewarded.

Sears ultimately did more than defend its intentions. It changed what the system rewarded.

Recalibrating Performance Measurement Required More Than Better Reporting

Sears announced several changes to its automotive compensation and oversight practices. Service advisors were moved from commission-based compensation to hourly pay, and product-specific sales goals were removed. The company introduced incentives connected to customer satisfaction and added independent, unannounced audits.

These changes did not make measurement disappear. They changed the balance of the measurement system.

Before the reforms, the strongest signals emphasized sales volume, target achievement, and selected product categories. Afterward, Sears attempted to place more weight on customer satisfaction, compliance, and independent review. The company did not merely change how it described performance. It changed the environment in which performance was produced.

That distinction matters. When a metric contributes to distorted behaviour, the solution is rarely another dashboard explaining the same number more clearly. The organization must reconsider the relationship between the metric, the incentive, and the outcome.

Removing commission reduced the direct financial reward attached to each additional sale. Eliminating product-specific goals weakened the pressure to recommend selected services. Independent audits created an external check on behaviour that sales data could not evaluate by itself.

Customer satisfaction introduced another perspective, although I would not treat it as a perfect replacement. Satisfaction scores can also be incomplete. Customers may be satisfied with unnecessary work because they lack the knowledge needed to assess it. They may be dissatisfied with necessary repairs because of cost. Employees can also learn to optimize satisfaction measures in ways that do not improve the underlying service.

The lesson is therefore not that Sears replaced a bad metric with a flawless one. The lesson is that sales alone could not safely represent the entire service objective. A stronger system required multiple signals and independent forms of verification.

This is what recalibration should mean: not the abandonment of measurement, but the recognition that no single operational metric should be allowed to define a complex outcome without challenge. I think this is where many organizations stop too early. They identify a problematic metric, replace it with another metric, and assume the interpretation problem has been solved. But a new number can become another incomplete proxy.

The stronger response is not simply to search for the perfect KPI. It is to build a measurement system capable of questioning its own signals. That may require quantitative data, qualitative evidence, independent review, customer feedback, operational audits, and clear limits on what any single metric is allowed to prove. The goal is not to eliminate uncertainty. It is to prevent one visible number from becoming the unquestioned definition of success.

Measure What Employees Can Influence—Then Examine How They Will Influence It

Organizations often rely on metrics because the outcomes they care about most are difficult to observe directly. It is easier to measure revenue than trust, speed than judgment, output than quality, activity than necessity, and short-term conversion than long-term customer interest. That does not make these metrics useless. It makes them incomplete.

The Sears case centered on automotive repair sales, but the interpretive problem it revealed is not limited to automotive service. In both that setting and many others, the same pattern applies: an organization finds a number it can observe, attaches compensation or performance pressure to it, and then interprets improvements in that number as evidence that the underlying outcome has improved.

This is the distinction organizations often miss in performance measurement. Management may view a KPI as a neutral representation of performance, while the employee reads it as an instruction about what to increase, protect, or prioritize.. That interpretive gap is what the Sears case makes visible — and it appears with equal clarity in digital performance measurement.

Click-through rate can rise while relevance weakens. Conversion rates can improve while customer quality declines. Content output can increase while usefulness deteriorates. A support team can close more tickets while leaving more problems unresolved. In each case, the number may be accurate. What becomes misleading is the assumption that improving the metric necessarily means improving the outcome it was chosen to represent.

A visible signal becomes convenient. Management connects it to a goal. Teams begin optimizing it. The number improves, and the organization assumes the underlying outcome improved with it. Sometimes that assumption is correct. Sometimes the metric has simply become easier to move than the outcome it was supposed to represent.

Before connecting a KPI to compensation, recognition, rankings, or employment decisions, an organization should therefore ask more than whether the performance measurement is technically accurate. It should ask what the metric genuinely represents, what important outcomes it omits, how people can increase it, and whether improving the number could weaken the broader result. It should also ask what balancing signals, qualitative evidence, or independent checks are needed to interpret the metric responsibly. Most importantly, it should examine how the measurement system will change the environment it is meant to observe.

This is why I do not think the central question is simply whether a number is correct. The more important question is what happens once people are rewarded for moving it.

A metric becomes dangerous when management treats it as a neutral window into performance while employees experience it as a command. The number may remain accurate. What changes is the behaviour organized around it.


Related Analysis