12 Check and Act: What Did We Actually Improve?
“Without data, you’re just another person with an opinion.” — W. Edwards Deming
12.1 Why This Chapter Matters
A countermeasure can be sensible, well designed, and faithfully carried out—and still fail to solve the customer problem. Check is where we compare that hope with what happened. Act is where we decide what the evidence earns: retain the method, revise it, abandon it, or begin another cycle around what remains. This chapter closes the guided Wash n’ Fold cycle before the capstone asks you to run one yourself.
12.2 Learning Objectives
By the end of this chapter, the student will be able to:
- Data Understanding Compare before-and-after process measures using consistent boundaries, cohorts, units, and calculation methods Bloom:Analyze
- Lean Six Sigma Principles and Tools Distinguish evidence from one matched observation, repeated descriptive evidence, and claims that require later statistical analysis Bloom:Evaluate
- Leading and Managing Projects Make an Act recommendation that names what to retain, revise, monitor, and investigate in the next PDCA cycle Bloom:Evaluate
- Business Understanding Present a bounded recommendation to a process owner using customer value, operating mechanism, business tradeoffs, evidence limits, and a clear decision request Bloom:Apply
Chapters 10 and 11 took us through Do: diagnose the observed waste, design a release rule, and create workplace conditions that make the proposed method executable. Now we will compare the proposed mechanism with after-state evidence. The aim is not to reward a favorable number. It is to decide what the evidence allows us to say.
12.3 Wash n’ Fold Project: Check
12.3.1 Return to the Trial Proposal Before Seeing the Results
Chapter 9 established the customer problem: six of 21 eligible orders missed 5:00 p.m. on the representative simulated Friday. The biggest measured delay was the release queue before Wash. Chapter 10 separated that evidence from attractive guesses; Chapter 11 turned the diagnosis into a flow-and-pull proposal supported by 5S and standard work.
Return to the Trial Proposal from Chapter 11 before reading the tested method below. Its mechanism, risks, measures, guardrails, and decision rules were written before the result was known. That order matters: Check evaluates a prior claim rather than inventing a favorable standard after seeing the answer.
A warning is warranted: 5S alone probably will not cut it. Reducing a few setup minutes can help, but it does not remove the rule holding the next batch until the active batch has cleared drying. A good proposal addresses that delay without simply flooding downstream work or hiding unprocessed orders outside the count.
12.3.2 The Method Tested in the Simulation
Compare the proposal with this tested package:
- Replace serial batch release with FIFO pull: release the next sorted order when a production-work slot is available.
- Limit released production WIP to 12 customer orders, returning a slot after production is complete.
- Request staging and notification as soon as an order is production-complete rather than waiting to assemble a notification batch.
- Put labeled supplies at the point of use, reducing Wash setup while retaining a small residual setup time.
This is a staged set of related changes, not the model for a capstone that tests one bounded countermeasure. Release control is the primary intervention; prompt notification and point-of-use supplies are supporting changes whose incremental evidence must be considered separately. The work limit is a modeled trial setting, not a universal recommendation. Orders waiting for a production slot still count as open customer orders. “Immediate notification” removes intentional batching; it does not promise instant service when both workers are occupied.
No workers, machine positions, operating hours, or nominal machine-cycle speeds were added. The matched scenarios use the same arrivals, order attributes, underlying service requirements, and correction decisions. This makes the comparison useful for examining the modeled rules, although it cannot establish what an actual laundry would achieve.
12.3.3 A Small After-State Evidence Packet
Begin with the same representative Friday as before: run 91, selected from the current state without inspecting its improvement. The after state is the future_pull_5s scenario for the same seed. There are still 39 orders, of which the same 21 qualify for the promise.
Chapter 9 reconstructed one complete order by hand so the software’s aggregation would be intelligible. Here, software performs that repetitive work for the after-state records. The learner’s job is to verify that the output uses the intended cohort, boundary, units, and denominator, then interpret the comparison.
The boundary is still drop-off to ready-and-notified. Minute zero is 7:00 a.m.; eligibility ends at minute 120; the promise deadline is minute 600. Closed hours are omitted from the cumulative operating clock.
| Order suffix | Arrival | Before operating L/T | After ready | After operating L/T |
|---|---|---|---|---|
| 001 | 9.1 | 165.2 | 143.7 | 134.6 |
| 002 | 10.9 | 162.8 | 152.8 | 141.9 |
| 003 | 12.1 | 166.3 | 172.1 | 160.0 |
| 004 | 31.8 | 254.5 | 225.6 | 193.8 |
| 005 | 33.7 | 253.2 | 211.8 | 178.1 |
| 006 | 44.1 | 246.4 | 221.8 | 177.7 |
| 007 | 47.8 | 343.1 | 229.9 | 182.1 |
| 008 | 52.8 | 339.1 | 291.4 | 238.6 |
| 009 | 60.3 | 334.3 | 292.3 | 232.0 |
| 010 | 61.2 | 424.5 | 334.8 | 273.6 |
| 011 | 62.1 | 423.2 | 295.4 | 233.3 |
| 012 | 63.6 | 425.2 | 330.3 | 266.7 |
| 013 | 75.3 | 518.1 | 362.8 | 287.5 |
| 014 | 78.2 | 511.5 | 366.9 | 288.7 |
| 015 | 81.0 | 509.5 | 340.5 | 259.5 |
| 016 | 88.0 | 598.3 | 372.8 | 284.8 |
| 017 | 89.7 | 596.0 | 386.3 | 296.6 |
| 018 | 97.4 | 591.4 | 412.8 | 315.4 |
| 019 | 100.0 | 682.0 | 408.8 | 308.8 |
| 020 | 103.1 | 678.4 | 426.0 | 322.9 |
| 021 | 110.5 | 676.2 | 450.5 | 340.0 |
The software-produced comparison for the 21 promise-eligible orders is:
| Eligible-order measure | Before | After |
|---|---|---|
| Mean operating lead time (min) | 423.8 | 243.6 |
| Median operating lead time (min) | 424.5 | 259.5 |
| Nearest-rank P90 operating lead time (min) | 676.2 | 315.4 |
| Orders ready after the 5:00 p.m. promise | 6 of 21 (28.6%) | 0 of 21 (0%) |
Complete Exercise 5, then write the provisional Check/Act recommendation in Exercise 6 before opening its answer. Confirm that the software used the same 21-order cohort and deadline as the baseline, and verify the after-state violation count directly from the ready timestamps. Use the supplied results for the remaining outcomes. The aim is to audit and interpret the calculation, not to repeat arithmetic already learned in Chapter 9.
| Supplied measure, all 39 orders | Before | After |
|---|---|---|
| Mean operating lead time (min) | 662.6 | 314.3 |
| PCE, total VA / total operating L/T | 13.9% | 29.2% |
| Correction occurrences | 5 | 5 |
These figures are derived from the displayed one-decimal occurrence timestamps. After-state VA time sums to 3,582.6 minutes and operating L/T to 12,256.6 minutes. The baseline VA sum was 3,582.3; that tiny difference is timestamp-rounding arithmetic, not a claim of changed service requirements. Keep this all-order readout separate from the 21-order promise cohort.
12.3.4 Check the Local Change and the Customer Result
The package changes both the release method and the supply arrangement, so examine it at two levels. For the supply trial, compare retrieval and setup under the current and proposed methods using the boundary defined in Chapter 11. Check dosing errors, replenishment interruptions, and handling or safety problems beside the timing. A faster setup that creates a new problem is not an improvement.
Then return to end-to-end lead time and eligible-order promise violations. Most of the baseline delay occurred while the release rule held the next batch, not while somebody measured detergent. A real setup saving could therefore leave the customer promise largely unchanged, while a release-rule change could improve the customer result even if setup barely moved.
The simulated packet reports the combined result of several changes and does not supply a separate observed setup-time series for the point-of-use trial. Do not use the package result to pretend that the supply arrangement alone caused the lead-time change. The proper component decision here is defer judgment and collect the missing setup evidence. In an actual pilot, retain the arrangement if the local measure improves without worse dosing, replenishment, or safety results; revise it when a correctable problem appears; abandon it when it produces no useful local change or creates a risk that cannot be acceptably controlled. Those judgments apply the decision rules written during Do.
12.3.5 What This Packet Does Not Establish
A better result on one matched Friday is evidence about that simulated Friday. It does not guarantee zero late orders next week, establish a real-world effect, or separate the contributions of four changes made together. The broader evidence in the answer to Exercise 6 follows the same 52 underlying Friday demand streams through each scenario. That gives us one simulated calendar year rather than a lucky afternoon, but it still comes from a model.
Repeated observations give us a better view of ordinary variation than one Friday can. They still do not tell us, by inspection alone, how much uncertainty surrounds an estimated improvement or how confidently an observed difference can be separated from ordinary variation. Later statistical methods will help us ask those questions rigorously.
Sometimes an intervention does not produce the result you expected. You may have chosen the wrong intervention. The intervention may have addressed one real problem while another constraint, defect, or source of delay remained beneath the surface. The process may already have been close to the best performance it could achieve under its current conditions, leaving little room for the proposed change to help. The method may also have been carried out inconsistently, or ordinary variation may have obscured a real but modest effect.
At White Belt, you do not yet have all the tools needed to decide confidently which explanation is most likely. You can still do disciplined work: verify what was actually implemented, compare consistent measures, report the unexpected result honestly, and identify the next observation or test that would separate the competing explanations. Later belts add deeper diagnostic, experimental, and statistical methods for making that judgment. Until then, do not force an appealing story onto an uncooperative result. Carry the uncertainty into the next PDCA cycle.
12.4 Wash n’ Fold Project: Act
12.4.1 Standardize What the Evidence Supports
An Act recommendation should say what you would retain, what you would monitor, and what would make you reconsider. In a simulation, “retain” means carry the method into the next modeled test or propose a bounded real pilot. Do not describe the result as an implemented gain in an actual business.
Assign an owner for the chosen release standard and supply locations. Review open customer orders as well as released production WIP, and track promise failures alongside setup, correction, and notification delays. A low count inside the production-work limit is no success if customers are merely waiting outside it.
State an escalation trigger before the next test: for example, the eligible-order violation rate exceeding the 10% goal over a defined review window, or repeated release-rule noncompliance. The window and denominator must be explicit; one small day with no failures cannot establish permanent reliability.
Compare this recommendation with the proposal you wrote before seeing the tested package. What did the evidence change? A well-reasoned alternative remains worth testing even if it differs from this worked example.
12.4.2 Presenting the Recommendation to the Owner
The following worked dialogue models a presentation to an owner. It shows how to request a trial, answer questions about the evidence, and incorporate a worker’s objection.
The project team does not open with the simulation model or the PCE calculation. It opens with the decision the owner has to make.
“We recommend a controlled trial of a pull release rule, prompt notification, and point-of-use supplies,” the presenter says. “The aim is to reduce missed 5:00 p.m. promises without increasing corrections or exhausting the staff. We are asking you to authorize a small real pilot, not to accept the modeled result as a guaranteed savings.”
The owner points at the row showing zero late eligible orders on Friday 91. “Then why not say the problem is solved?”
“Because that was one simulated Friday. Across 52 modeled Fridays, the complete package still produced five late orders. The result is encouraging, and the matched comparisons support the mechanism, but no customer order has moved through your shop under this rule yet.”
The owner asks which part of the package did the work. The team shows the incremental scenario comparison. Release control produced the largest modeled reduction in operating lead time; prompt notification and point-of-use supplies added smaller gains. The team also states what the simulation did not improve: correction decisions were held fixed, so the package supplies no evidence that the tagging and relabeling problem has been solved.
“Could we get the same result by buying another washer?” the owner asks.
“Possibly, but the current evidence does not identify washer capacity as the dominant cause of the missed promise. The largest observed delay came from the rule that held work before Wash. Buying capacity before testing the release rule would cost more and might leave that delay intact.”
One of the workers then raises a practical objection. Twelve released orders may be manageable on an ordinary Friday, but a few unusually large orders could crowd the folding stations and make the limit unsafe or misleading. The team adds order size and blocked-work visibility to the trial checks and agrees that the workers can stop the release when the defined conditions are not met. The modeled limit of 12 becomes a starting hypothesis to examine, not a command imported from a spreadsheet.
The owner authorizes planning for a bounded pilot, with the final release limit set only after the workers review recent order sizes and available space. One person will own the release signal and another will verify the ready-and-notified timestamps. The review will compare late-promise rate, operating lead time, all open orders, correction occurrences, and worker-reported overload using the same definitions as the baseline. If quality worsens, work becomes unsafe, or the release rule repeatedly breaks down, the team will stop and revise the trial.
A compact owner brief should state:
- Decision requested: what authority, resources, or trial approval is needed now.
- Customer and business problem: who is affected and why the result matters to the firm.
- Proposed mechanism: what will change in the work and why it could affect the problem.
- Expected benefit and burden: which outcomes may improve, what effort or cost the change requires, and who carries it.
- Evidence and limits: what was observed, calculated, or simulated and what remains unknown.
- Measures and review rule: what will be monitored, who owns it, when the result will be reviewed, and what would stop or revise the trial.
The owner’s questions help validate the recommendation. A recommendation becomes more credible when it survives questions about cost, feasibility, worker burden, and the difference between a modeled result and an implemented one.
12.5 Exercises
1. Bloom:Evaluate Suppose your pilot shows improved lead time but a slight rise in defect/rework rate. Would you standardize now, revise and re-test, or roll back? Defend your decision.
Most strong responses choose revise and re-test unless defect risk is trivial and controlled. White Belt judgment should balance flow gains against quality stability before standardization.
2. Bloom:Apply Write a one-paragraph Act-phase standardization note for one improvement that worked in your case. Include owner, trigger, review cadence, and the condition that would reopen the method for revision.
A good note names who owns the new standard, when it is triggered, how often compliance and performance are reviewed, and what result requires reconsideration. It should be specific enough that another team member can execute it without ambiguity.
3. Bloom:Apply Draft the opening Plan statement for the next PDCA cycle based on what remained unresolved. Include one problem statement sentence and one SMART goal sentence.
Answers vary. A strong response is concrete, measurable, and explicitly connected to observed results from the prior cycle.
4. Bloom:Evaluate Your after-state improves lead time but leaves rework unchanged. Would you accept, revise, or reject the change? Defend your decision using two criteria.
The evidence can justify retaining the lead-time change while revising the next cycle to address rework; it does not justify claiming that the whole problem is solved. A strong response uses at least one customer or flow criterion and one quality or risk criterion, checks any predeclared guardrails, and explains why the chosen Act decision fits both. Reject or roll back the change if the unchanged rework violates an acceptance rule or leaves an unacceptable harm; otherwise preserve the supported gain and keep the correction gap visible.
5. Bloom:Analyze Audit the software-produced 21-order comparison. Confirm that its cohort, boundary, clock, physical unit, percentile convention, and promise denominator match the baseline. Verify the after-state violation count directly from the ready timestamps, then write a short Check summary that also notes the supplied all-order PCE result and unchanged correction count.
The comparison uses the same 21 eligible orders, drop-off-to-ready boundary, operating-minute clock, minutes-per-order unit, nearest-rank percentile convention, and 5:00 p.m. deadline as the baseline. All 21 after-state ready timestamps are at or before minute 600, confirming the software’s zero violations, compared with six of 21 before. Mean operating lead time falls from 423.8 to 243.6 minutes, and P90 falls from 676.2 to 315.4 minutes. The zero violation result applies to this matched Friday, not every future Friday.
Across all 39 orders, the supplied PCE rises from 13.9% to 29.2%, while correction remains five occurrences. The package improves modeled flow and promise performance but does not remove the correction problem. Do not mix the all-order mean of 314.3 with the eligible-order mean of 243.6, or use all 39 orders as the promise denominator.
6. Bloom:Evaluate Before opening the answer, make a provisional Check/Act recommendation from this Friday. What would you retain for another test, what would you monitor, and what can you not yet attribute to 5S? Then inspect the calendar-year simulation evidence in the answer and explain whether it changes your recommendation.
A defensible provisional recommendation is to retain the package for further testing, monitor promise violations and all open orders, and continue investigating correction. Do not attribute the whole improvement to 5S: release, notification, and setup changed together. An alternative proposal is not wrong merely because it differs from this package; it needs comparable evidence.
The simulation also follows 52 calendar Friday demand streams through every scenario. Each matched set uses the same arrivals and underlying work requirements; the scenarios differ only in their specified operating rules. The Friday arrival intensity follows an explicitly modeled annual cycle ranging from 15% below to 15% above the typical level. This is an illustrative assumption for testing the methods under quieter and busier periods, not a claim about the seasonal pattern of an actual laundromat. The table reports averages of the 52 Friday-level measures, not one pooled set of orders.
| Scenario | Mean operating L/T (min) | Mean daily P90 (min) | Mean daily miss rate | Mean setup/order (min) |
|---|---|---|---|---|
| Current serial batches | 657.6 | 1,053.5 | 27.60% | 4.23 |
| Pull only; batched notification; shared jug | 365.0 | 494.6 | 1.66% | 4.23 |
| Pull + prompt notification; shared jug | 339.8 | 468.8 | 0.43% | 4.23 |
| Pull + prompt notification + point-of-use supplies | 324.2 | 452.4 | 0.31% | 2.33 |
Release control accounts for the largest incremental reduction in mean operating L/T: approximately 292.6 minutes from current batches to pull only. Prompt notification adds about 25.2 minutes; point-of-use setup adds about 15.6 beyond that. These are differences between scenario means in this sequence, not independent effects guaranteed under every combination. There is no 5S-only scenario here, and no basis for claiming that setup changes alone would reproduce the combined gain.
Pull-only mean operating L/T is lower than current-state mean operating L/T on all 52 matched Fridays. That consistency supports the modeled release explanation; it is not independent validation of the model or a forecast for a real shop. The final scenario still has five late eligible orders across the simulated year, so the zero failures on Friday 91 should not be generalized.
Two aggregation details matter. First, these technical exports calculate each day’s P90 by interpolation, then average those daily percentiles; they are neither the pooled P90 nor the nearest-rank P90 reported for the matched Friday in Exercise 5. Second, the mean daily violation rates weight Fridays equally. Pooling orders gives \(314/1{,}007 \approx 31.18\%\) before and \(5/1{,}007 \approx 0.50\%\) after, because Fridays have different eligible-order counts. Both summaries are valid when labeled; they answer different questions.
The stronger Act recommendation is to carry the release change forward with supporting notification and setup standards, then test whether an actual shop can implement them safely and reliably. Keep a named owner, a defined review window, a promise denominator, and an escalation trigger. Monitor correction separately: the simulation held correction decisions fixed, so improved flow does not demonstrate improved defect prevention.
7. Bloom:Apply Prepare a six-bullet owner brief for your proposed improvement using the categories above. End with one difficult question you would want the owner or a process participant to ask before authorizing the trial.
Answers vary. A strong brief requests a bounded decision, connects the problem to customer and business consequences, explains the proposed mechanism, names costs and human burdens, distinguishes observed or simulated evidence from unknowns, and defines ownership and a review rule. The closing question should expose a real vulnerability—for example, whether the work limit fits large orders, whether another measure could worsen, or whether the people performing the work can carry out the method safely.
12.5.1 Journal and Reflect
Which part of your original proposal changed after you inspected the evidence? What result did you find most persuasive, and what uncertainty still limits the recommendation?
Experimentation and management both require some tolerance for ambiguity. Describe a time when you had to make a responsible decision before the evidence clearly favored one explanation. How could you remain open to several plausible explanations without allowing uncertainty to become indecision? Name one practice that would help you take a concrete next step without pretending to know more than you do.
The guided cycle is now complete. The capstone asks you to reconstruct the same reasoning with a process of your own.
12.6 Chapter Summary
- Check compares the after state with the baseline using the same boundary, cohort, units, and calculation method.
- One favorable result supports a provisional judgment; repeated comparable observations show whether the pattern survives ordinary variation.
- Descriptive evidence does not by itself quantify uncertainty or prove that an intervention caused a real-world effect; later statistical methods strengthen those judgments.
- Act states what to retain, revise, monitor, and investigate, then carries the remaining gap into the next PDCA cycle.
- A useful owner presentation connects customer value, operating mechanism, business tradeoffs, evidence limits, and a specific decision request.