Our Services

Where we work

We work with leadership teams whose operations have plateaued — applying the engineering behind Pull, Variability, and Flow to see where the familiar Lean toolkit stops, and where AI begins to change the operating model itself.

01

Operations Diagnostics

An honest assessment of where your current system is delivering, where it has plateaued, and why the metrics may be hiding the gap.

02

Demand Driven Advisory

Guidance on DDMRP, DD Operating Model, and Demand Driven S&OP — moving from forecast-bound planning to flow under real demand.

03

AI-Readiness for Operations

Cutting through the noise: where AI genuinely changes the operating architecture, and where it is merely cost without leverage.

04

Executive Workshops

Working sessions for leadership teams — building the shared language and the strategic conversation that change requires.

Insights

Selected thinking

Essays on operations, the limits of Lean, and the shift to an AI-native way of working. Written when it matters, not on a schedule.

April 2026

The four traps that keep capable plants locked in place

How tools designed to surface problems quietly become the scorecards that hide them.

April 2026

AI is not a tool to bolt onto Lean

The mistake most operations leaders are about to make — and the architecture question underneath it.

May 2026

Why your OEE is rising while you fall further behind

Absolute improvement and relative decline can happen at the same time — and most metrics are built to miss it.

May 2026

What the West imported from Toyota — and what it left behind.

Lean delivered real gains for a generation. Here is the part of the system that never made the trip.

June 2026

Buffers are not waste.

Why a generation of inventory orthodoxy got one of the most important questions in operations exactly backward.

July 2026

Why smart executives keep funding the wrong fix.

The problem was never a shortage of intelligence. It is that the wrong fix feels responsible and the right one feels reckless.

August 2026

Asking AI the right question about your operation.

A large language model is built to agree with you. Everything useful it can do for your operation depends on asking it the one question that forces it not to.

September 2026

Lean + AI won't work.

The problem isn't the AI. It's the plus sign — and what you are being told to add it to.

October 2026

What replaces forecasting.

Everyone knows the forecast is wrong. The strange part is that the entire planning system is still built to obey it.

September 2026

Lesson 1 — The question that changes the conversation.

A $650M food manufacturer spent $2.8M improving its forecast. MAPE stayed at 60%. The problem was not the forecast. It was the question.

October 2026

Lesson 2 — When the system that passed every audit becomes the problem.

A $420M medical device manufacturer passed every FDA audit for a decade. Then the CAPA backlog doubled, and the architecture that kept the plant compliant became the architecture that kept it behind.

Contact Us

Start a conversation

Tell us about your operation and what you are trying to change. We read every message.

Villeda Consulting Group, Inc.

Based in the United States · Engagements worldwide

beyondlean (at) villedaconsulting.com

The Book

Beyond Lean

From Toolkit to Thinking System — Recovering TPS in the AI Age

The book

Lean got you this far. It won’t get you through the AI age.

Beyond Lean argues that the West imported the visible tools of the Toyota Production System — 5S boards, SMED, kanban cards, kaizen events, Gemba Walks — and left the engineering thinking behind. The result is a toolkit that produced real gains for a generation, then quietly stopped moving the needle. What went missing was the engineering behind the tools — Pull, Variability, and Flow, designed together rather than borrowed as artifacts — and the multi-echelon architecture that Toyota spent three decades extending from the plant floor outward through its supplier network. The West imported the first stage and never built the rest. What replaces the toolkit is not another tool. It is a thinking system.

Most operations have plateaued. The temptation is to bolt AI onto the existing Lean program and expect a step change. It won’t come from addition. The exit is two tracks, run at once: a tourniquet to stabilize what’s bleeding inside the current architecture, and a surgery to rebuild what’s broken underneath it. The book explains why bolting AI onto Lean skips both and names the precise specialties — Pull, Variability, and Flow — that the surgery has to put back on the table.

The book moves in five parts. Part I traces what TPS actually was, what the West imported, and what was quietly discarded on the way. Part II exposes the Lean Paradox — the mechanism by which programs designed to surface problems become the scorecards that hide them — and names the business delusions that keep it running long after leaders have privately concluded the tools have stopped working. Part III maps the twelve biases and four traps that keep capable operations locked in place. Part IV shows why bolting AI onto Lean will not close the gap, and lays out the practical method for using an LLM as a thinking partner instead. Part V takes the argument to the C-suite: five decisions in sequence — Protect, Stop, Stabilize, Redesign, Build — and the burden-of-proof discipline that determines whether the COO defers or acts.

Beyond Lean is written for the manufacturing/operations VPs, plant managers, engineering, supply chain, and continuous improvement leaders — and for the CEOs and CFOs who fund their programs — who want the engineering and AI as a thinking partner rather than slogans. It is the conversation the Lean program was never built to have.

Get the Book

Beyond Lean: From Toolkit to Thinking System — Recovering TPS in the AI Age

Available in paperback ($38.95) and Kindle ($9.99).

Hosting a reading group? A five-session companion guide is available. Contact us with subject: reading group (at) villedaconsulting.com.


*A Kindle edition is planned.

About the Author

Ramiro Villeda, Ph.D., helped bring Lean to the West and spent thirty years implementing it across five continents. In 1985 he defended the first doctoral dissertation on the engineering of JIT pull systems in the Americas, under the mentorship of Shigeo Shingo, a principal architect of the Toyota Production System.

He is the founder of Villeda Consulting Group, Inc. He advises leadership teams on the engineering behind Pull, Variability, and Flow — and on where AI genuinely changes the operating model.

For permissions and consulting inquiries: Get in touch. Or go back.


The four traps that keep capable plants locked in place

How tools designed to surface problems quietly become the scorecards that hide them.

April 2026


Walk into a plant that has been "doing Lean" for fifteen years and you will usually find good people working hard at a system that is not improving. The value-stream maps are on the wall. The 5S audits are scored. Kaizen events happen on schedule. Standard work is documented. By every internal measure the organization is disciplined and engaged — and by every measure that matters to a customer, it has plateaued.

The reason is not a lack of effort or talent. It is that the tools meant to reveal problems have, over time, become the rituals that conceal them. This happens in a predictable sequence — four traps that feed one another until motion is mistaken for progress and the whole thing locks in place.

The Silo Trap. Improvement is mostly organized by department, because that is how the org chart is drawn and how budgets are defended. Each area optimizes itself: the stamping cell raises its utilization, the weld line cuts its changeover, purchasing lowers unit cost. Every silo improves, and the flow between them — where the customer's lead time actually lives — gets worse. The plant becomes a collection of locally excellent operations that collectively cannot deliver. Nobody owns the space between the boxes, so nobody sees the problem that lives there.

The Frozen Map Trap. At some point the organization drew a map — of the value stream, the layout, the standard process — and the map was good. Then the map stopped moving while reality kept going. Demand shifted, the product mix changed, a supplier moved, variability rose. But the map had become the official version of the truth, and questioning it felt like questioning the improvement program itself. So the plant keeps steering by a chart of a coastline that has since eroded. The tool that once captured reality now protects the organization from having to look at it again.

The Checklist Trap. To make improvement repeatable, it gets turned into a checklist — the 5S audit, the standard-work sheet, the leader standard work routine, the tiered accountability board. This is a reasonable instinct; checklists are how you keep good practice from decaying. But a checklist answers "did we do the steps," not "did the steps still make sense." Compliance becomes the goal. The audit score goes up while the thing the audit was supposed to guarantee goes down, because everyone is optimizing the checkmark rather than the outcome it was meant to stand for. The plant gets very good at passing its own tests.

The Stagnation and Blame Trap. When results stall despite all this visible discipline, the organization needs an explanation — and the easiest one to reach for is people. The improvement isn't landing because the operators aren't engaged, the supervisor isn't holding the line, the last kaizen didn't stick. Attention turns to individual accountability precisely when the problem is architectural. Blame is comfortable because it is local and actionable; redesigning the system is neither. So the plant runs more events, tightens more audits, and looks harder at the people — which returns the organization, stronger, to the Silo Trap it started in.

That is the loop. Each trap makes the next one feel like the responsible next step. And every trap shares a single property that makes it so durable: it cannot be seen from inside itself. From within the Silo Trap, optimizing your own department looks like discipline. From within the Checklist Trap, a rising audit score looks like progress. The traps are not failures of intelligence or will. They are what happens when tools built to expose reality are gradually repurposed into scorecards that conceal it.

Breaking the cycle does not start with a better tool. It starts with the uncomfortable recognition that the plant's own instruments have stopped telling it the truth — and that the busyness on the dashboard is not the same as improvement. Capable plants do not stay stuck because they lack capability. They stay stuck because every trap rewards the behavior that deepens it. The way out is not another tool or a harder push on the current ones — it is stepping outside the system long enough to see what the scorecards can no longer show you, and rebuilding the metrics so they reward flow instead of motion. That is difficult precisely because it cannot be done from inside the loop. But it is the only move that changes the trajectory rather than the speed.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He helps operations leaders see the traps their own scorecards are hiding. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


AI is not a tool to bolt onto Lean

The mistake most operations leaders are about to make — and the architecture question underneath it.

April 2026


Here is the plan most manufacturers are quietly converging on: take the Lean program you already run, and add AI to it. A copilot for the planners. A vision system on the inspection station. A demand-forecasting engine feeding the same S&OP process you have now. Digitize the toolkit, bolt on the intelligence, keep everything else where it is.

It is the safe-looking move, and it is the wrong one — for the same reason that digitizing a bad process gives you a faster bad process. AI's value in operations lies less in automating the tasks you already do than in enabling new architectures of coordination and decision-making. Bolt it onto an unchanged architecture and you have bought motion, not progress: the same decisions, made a little faster, inside a system that was already the constraint.

The instinct to bolt on comes from treating AI as another tool in the kit — the newest entry in a lineup that runs kanban, 5S, SMED, VSMs, and now a language model. But the tools were never the point, and neither is this one. What matters is the structure they operate inside: how work is coordinated across the plant, how information moves, how decisions get made and by whom. That structure is the thing a generation of toolkit Lean quietly left unexamined, and it is exactly the thing AI is capable of changing — if you let it.

Which surfaces the question almost no one is asking. Not "where can we apply AI to what we already do," but:

If you were designing your operation today — for your actual product mix, your actual demand variability, your actual supplier ecosystem, with these capabilities available from the start — what would you build?

That is the architecture question. It is uncomfortable because it does not have a tool-shaped answer. You cannot buy it, pilot it in one cell, or add it to next year's kaizen calendar. It asks you to reconsider the system, not to shop for a component.

The organizations that will pull ahead are not the ones that adopt AI earliest. Adoption is easy; everyone will do it, and a capability everyone has confers no advantage. The advantage goes to whoever redesigns the coordination and decision-making architecture around what these tools now make possible — and is willing to let that redesign change the org chart, the metrics, and the planning process, not just the software license.

This is not a manufacturing-specific observation. Strategists watching AI reshape the broader knowledge economy have arrived at the same conclusion from the opposite end of it: the durable gains come from redesigning how a system coordinates and decides, not from automating its individual tasks. The factory floor is simply where the lesson is most concrete, because a factory should be a coordination system you can walk through.

So before the copilots get provisioned and the forecasting engine gets a budget line, the honest first question is not "what can we automate." It is whether the architecture those tools are about to be bolted onto is the one you would design today — or is the one you inherited and stopped questioning a long time ago. Bolt AI onto the second, and you will get a faster version of a plant that had already plateaued. The tool will work exactly as advertised. The system will not move.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He works with operations leaders on the architecture question underneath the AI decision. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


Why your OEE is rising while you fall further behind

Absolute improvement and relative decline can happen at the same time — and most metrics are built to miss it.

May 2026


Every quarter the dashboard is greener. Overall Equipment Effectiveness is up another two or three points. Availability, performance, and quality have all moved in the right direction. The continuous-improvement program is working — the numbers say so. And yet the things that actually matter to a customer, on-time delivery and lead time, have not improved, or have quietly gotten worse.

This is not a paradox. It is the predictable result of measuring the wrong thing well.

OEE measures how hard a piece of equipment works. Multiply availability by performance by quality and you get a clean percentage that tells you a machine was running, running fast, and running clean. That is a real and useful thing to know about a machine. It is not the same thing as delivering more of what the customer ordered, when they ordered it — and treating it as if it were is where capable plants quietly lose ground.

A factory is not a collection of machines to be maximized one at a time. It is a system with a constraint. You know this already: in most plants, one or two resources set the pace of the whole operation, and everything upstream and downstream of them is either feeding them or draining them. Every hour you spend lifting the OEE of a resource that is not one of those pace-setters is an hour spent producing inventory the constraint cannot yet absorb. The dashboard turns green. The floor fills with work-in-process. Lead time lengthens. Cash locks up between operations. And OEE, blind to every bit of it, keeps climbing.

That is the mechanism behind the divergence. You are improving — genuinely, measurably — against a yardstick that has stopped predicting competitiveness. Meanwhile the competitor who measured flow, and left their non-constraints deliberately underutilized, shipped faster and pulled away. Both plants improved. Only one of them improved at something that mattered.

The deeper problem is behavioral. What you put at the top of the monthly operations review — the number the plant manager answers for — is a set of instructions to the organization about what to protect. Tell a plant that its scorecard is equipment utilization, and it will run every machine it can, because idle equipment looks like waste and green numbers look like success. That behavior is perfectly rational under the metric and quietly destructive to the system. A rising OEE at a non-constraint is not a sign of health. More often it is the fingerprint of overproduction — the oldest waste in the Toyota canon, reappearing in the disguise of a good number.

So what should sit at the top of that review instead? Metrics that describe the system rather than the asset:

Schedule adherence — did you build what you said you would, in the sequence you said — belongs in the conversation too, but it is an older metric, and on its own it will still reward a plant for faithfully executing the wrong plan. That is the tell that the real problem is bigger than the scorecard. What these flow measures actually point to is not just a better list of numbers but a different way of running the plant: a flow-based operating model, built on Demand Driven principles, in which the operation is coordinated around the movement of real demand through the system rather than the local utilization of its assets. The metrics are downstream of that choice. Adopt the model and the right numbers follow; keep optimizing assets and no dashboard will save you.

Notice what these have in common. Every one of them is a statement about the system's ability to convert a customer order into a shipment. Not one of them can be improved by running a non-constraint harder. That single property — that you cannot game them by keeping non-constraint machines busy — is what makes them worth measuring.

None of this makes OEE useless. At the constraint, OEE is exactly the right question: that is the one place where an hour lost is an hour lost for the entire plant, and squeezing every point of availability and speed and yield out of it is pure gain. The error is not measuring OEE. The error is applying an asset-level metric to a system-level problem — everywhere, on every line, on every machine, as the primary scorecard — and then being surprised when a plant full of green numbers keeps losing orders.

The metric you elevate is the behavior you will get. Choose one that rewards local busyness, and you will build a plant that is busy, green, and falling behind. Choose one that rewards flow, and the improvement finally appears where the customer can feel it — which is the only place it was ever supposed to appear.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He helps manufacturers replace local-efficiency scorecards with flow-based operating metrics. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


What the West imported from Toyota — and what it left behind

Lean delivered real gains for a generation. Here is the part of the system that never made the trip.

April 2026


Lean worked. For thirty years, Western manufacturers that adopted it cut inventory, shortened changeovers, cleaned up their floors, and eliminated waste that had been invisible to them for decades. That is not in dispute, and this is not an indictment of Lean.

It is a case about what got copied and what got left behind.

When the West studied the Toyota Production System, it brought home the things it could certify and teach: 5S boards, SMED, kanban cards, kaizen events, Gemba Walks, value-stream maps. What it left behind was the engineering reasoning that turns a one-time cleanup into a compounding system.

Toyota's engineers understood buffers as engineered decouplers — inventory placed deliberately, in calculated amounts, at specific points, to absorb variability and protect flow. The toolkit imported a slogan instead: inventory is waste, drive it to zero.

Toyota's engineers treated the number of kanban cards as a control parameter, calculated from demand variability, process variability, and changeover frequency. The toolkit imported the cards as a visual technique and the instruction to reduce them. The card became a ceremony; the math stayed in Japan.

Toyota's engineers knew that a perfectly balanced line is rarely the most efficient one under variability, and that deliberate imbalance can outperform it. The toolkit imported "balance the line" as a universal rule, then blamed the resulting fragility on execution.

Toyota's engineers treated standard work as a testable hypothesis — the best-known method so far, meant to be challenged and improved. The toolkit imported it as a compliance document: follow the standard, pass the audit.

And underneath all of it, Toyota's engineers treated the people on the floor as the sensors of the system. The toolkit imported the suggestion box and left the sensors behind.

None of this makes the gains fake. Any operation running for decades without disciplined waste elimination has accumulated enormous amounts of obvious waste — and the visible tools are highly effective at surfacing it. That part doesn't require the engineering underneath. It just requires the discipline to look. What it doesn't do is keep working once the obvious waste is gone, because the part of the system built to generate the next architectural insight was exactly the part that never crossed the ocean.

That is why so many mature Lean programs plateau at the same ceiling, a decade or more in: audits passing, boards green, nothing left to squeeze. Recovering the difference is not a matter of running more events. It is going back for the engineering reasoning — about flow, variability, buffers, and pull — that the tools were always meant to express, not replace.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. His work is recovering the engineering thinking beneath the Lean toolkit. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


Buffers are not waste

Why a generation of inventory orthodoxy got one of the most important questions in operations exactly backward.

June 2026


Somewhere in the translation of the Toyota Production System into Western practice, a piece of engineering became a commandment. Inventory is waste. Drive it to zero. It is printed on training slides, scored on audits, and repeated on plant floors as if it were a law of nature. It is also, as stated, wrong — and the cost of believing it has been enormous.

The engineers who built TPS did not believe inventory was waste. They believed inventory was information — a visible signal of variability, imbalance, and unreliability in the system. Their goal was never zero inventory for its own sake. It was to expose the problems that inventory was hiding, fix them, and then remove the inventory that was no longer needed. Reducing inventory was the consequence of eliminating variability, not a shortcut around it. The Western toolkit inverted this. It kept the target — less inventory — and discarded the reasoning, which is how a diagnostic instrument became a blunt directive to cut the one thing that was holding together a system subjected to variability.

Because here is what a buffer actually is when a supply chain specialist or an engineer designs it: a decoupler. A deliberate, calculated quantity of inventory, time, or capacity, placed at a specific point, whose job is to absorb variability so that a disruption in one part of the system does not immediately starve or stall the next. Remove the buffer without first removing the variability it was absorbing, and you have not eliminated waste. You have removed the shock absorber and kept the potholes. The line is now "leaner" and dramatically more fragile, and the first mechanical breakdown, supply hiccup or demand spike propagates straight through it.

This is the distinction the orthodoxy erased: the difference between a buffer that is engineered and inventory that has merely accumulated. Accumulated inventory — piles of work-in-process sitting between operations because the system is unbalanced and no one decided otherwise — genuinely is waste, and it is what the "inventory is waste" slogan was originally reacting against. An engineered buffer is the opposite: a designed element of the system, sized from the actual variability of demand and process, positioned where it protects flow. One is a symptom. The other is a control. Treating them as the same thing, and attacking both with the same zeal, is the central error.

The right question was never "how do we get to zero." It is "where does this system need to be decoupled, by how much, and why." That question has real answers — they depend on demand variability, process variability, changeover frequency, and where the constraint sits — and the answers are calculated, not chanted. A plant that asks it ends up with less inventory than one drowning in accumulated work-in-process, but more than one that stripped its buffers to hit a slogan and now expedites daily to survive. The engineered answer sits in the middle, and it moves as conditions move.

None of this is an argument for carrying more inventory. It is an argument for carrying the right inventory, in the right place, for a reason you can defend — and for recognizing that a buffer sized to protect a customer commitment is not a failure of discipline. It is discipline. The failure is the unthinking one: seeing every buffer as waste, cutting it to improve a number, and mistaking the resulting fragility for "efficiency" right up until the day the system can no longer absorb a shock it used to shrug off.

A generation of operations leaders was taught to view inventory as the enemy. It has always been the messenger. Shoot the messenger and you do not solve the problem it was reporting — you just lose the ability to see it coming.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He helps manufacturers distinguish engineered buffers from accumulated waste — and size the difference. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


Why smart executives keep funding the wrong fix

The problem was never a shortage of intelligence. It is that the wrong fix feels responsible and the right one feels reckless.

July 2026


Sit in enough operations reviews and you notice a pattern that has nothing to do with competence. A capable, experienced executive — someone who has run plants, hit numbers, earned the corner office — looks at a program that has plateaued and decides to double down on it. Another Lean rollout. A fresh wave of kaizen events. A digital layer bolted onto the same process. The gap does not close, and a year later the same executive funds the next version of the same fix.

This is not stupidity. These are smart people. That is exactly why it happens.

The wrong fix survives because several forces press on experienced decision-makers at once, and all of them point the same way.

The halo of past success. When a Lean program delivered real gains a decade or two ago, the credit attached to the tools — the boards, the cards, the events — and it has never come off. So when results stall, "more of what worked" feels obviously right. What the halo hides is that those early gains came from picking low-hanging fruit that any disciplined effort would have captured. The credit went to the specific toolkit anyway, and the toolkit became above suspicion. You cannot fix what you are not allowed to doubt.

The answer already within reach. Under strong time pressure, the mind reaches for the available answer, not the correct one. Rethinking the architecture of an operation is slow, effortful, and uncertain. Running another event is fast, familiar, and easy to defend in the meeting. Against a quarterly clock, the fast answer wins — not because the executive weighed it and judged it better, but because it arrived first and cost less to think.

The illusion of understanding. Organizations believe they understand the Toyota Production System because they can see its tools sitting on the floor. But visibility is not understanding. The engineering logic beneath those tools — the reasoning about pull, variability, buffers, and flow — was never imported with them. So the organization is confidently operating a system it does not actually understand — and there is nothing harder to dislodge than confidence in a model no one realizes is incomplete.

Blame is easier than redesign. When the program underdelivers, attention turns to people — the plant did not execute, the team was not engaged, the last kaizen did not stick. Blaming execution is comfortable because it is local and actionable. Questioning the architecture is neither. So the organization keeps its model and changes its people, which quietly guarantees the next version fails in the same way as the last.

The system pulls back. Every incentive, every metric, every reporting line was built around the current architecture. Change threatens all of them at once. An organization settles into a stable pattern the way a marble settles into a bowl, and the sides of the bowl pull every decision back toward the familiar — even when everyone can see the familiar is no longer working.

Put these together and the picture is not a smart executive making a dumb decision. It is a smart executive making the decision the situation is built to produce. The wrong fix is cognitively cheaper, socially safer, blessed by past success, and protected by every structure in the building. The right fix — stepping outside the model and questioning the architecture itself — is slow, threatening, and unrewarded until the day it finally works.

Which is why the pattern almost never breaks from the inside. The forces that sustain it press hardest on the people most invested in the current system — which is to say, the very people with the authority to change it. Those with the least at stake can often see the shape of the bowl from the rim; the executive who built its walls sits at the very bottom, where the curvature is invisible and each pull back toward the familiar registers not as a constraint but as sound judgment.

Breaking out, then, is rarely a matter of more resolve. It takes one of two things: an outside vantage point with no stake in the existing answer, or a discipline that forces the architectural question onto the table before the fast, familiar answer can crowd it out. The goal is not to try harder inside the model. It is to get the real question asked at all.

The executives funding the wrong fix are not failing to think. They are thinking exactly as their situation rewards. That is what makes the trap so durable — and why escaping it begins not with working harder, but with noticing that the smart-feeling move and the genuinely smart move have quietly come apart.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He helps leadership teams surface the architectural questions their own incentives keep off the table. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


Asking AI the right question about your operation

A large language model is built to agree with you. Everything useful it can do for your operation depends on asking it the one question that forces it not to.

August 2026


The large language models (LLMs) arriving in operations are genuinely capable. The trouble is not the tool. It is that most leaders will aim it at the wrong question — and a powerful tool aimed at the wrong question returns a confident, useless answer faster than anything before it.

The wrong questions are the ones already on the table. "Where can we automate?" "Optimize this schedule." "Cut the changeover on line 3." These are task questions, and the model will happily oblige — it will optimize the schedule inside the plant you already have and automate the task inside the process you already run. What you get back is a faster version of the operation you already plateaued with.

There is a deeper problem, and it is specific to how these tools work: an LLM is built to be agreeable. Ask whether your current approach is sound and it will find reasons you are right. Frame a question around your existing setup and it will improve things within that frame and never challenge it — because you never asked it to. Left to its defaults, it hands your own assumptions back to you with better polish. The skill here is not prompting. It is refusing to let the machine agree with you.

Which points to the one question worth asking. Not "how do I improve what I have," but this:

What would we build if we started today, with no legacy and AI as the foundation?

The phrase doing the work is "no legacy." It forces the model off the agreeable path and onto the architecture, because it can no longer default to improving what you already have. It has to describe something built clean, then let you see the distance between that and what you are running now.

To make that question do its work, two disciplines:

Give it the real context first. Where the constraint sits, where variability actually lives, where cash and lead time are trapped — before you ask anything. An LLM reasoning from your specifics gives sharper answers than one reasoning from generic best practice, and it is far harder to wave away.

Then make the LLM argue against you. Ask it to build the strongest possible case that your current architecture is the wrong one. Ask what a competitor designing from a blank sheet would do differently. Ask what it would remove.

The value isn't the disagreement. It's watching your own data trace back to the mechanism producing the gap.

None of this is about surrendering judgment to a model. It is about using the model to surface the questions your own organization is structured not to ask — the architectural ones that every incentive and reporting line quietly keeps off the table. The LLM is not a thinking partner because it is clever. It is a thinking partner because it has no stake in your current system and no attachment to the answer that keeps the peace — but only if you ask it in a way that puts that detachment to use.

So before the pilots and the copilots and the budget lines, the first question is not what the model can do for your operation. It is whether you are willing to ask it the one question your operation was built to avoid — and to insist it answer you honestly, rather than agreeably.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He helps operations leaders turn AI from an agreeable assistant into an honest thinking partner. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


Lean + AI won't work

The problem isn't the AI. It's the plus sign — and what you are being told to add it to.

September 2026


Walk the floor of any operations conference this year and you will hear the same equation, over and over, from vendors, consultants, and keynote stages: take your Lean program, add artificial intelligence, and watch the transformation follow. Lean plus AI. It is the formula of the moment.

It will not work — and not because AI is overhyped. AI is genuinely powerful. It will not work because of what the plus sign quietly assumes, and what you are being told to add it to.

The plus sign assumes AI is an addition: one more capable tool, laid on top of a system that already works. But the system underneath does not work the way you think it does. Western Lean is Toyota's approach with the engineering thinking stripped out. It kept the visible practices — the boards, the cards, the value-stream maps — and left behind the reasoning about pull, variability, buffers, and flow that made those practices mean anything. So "Lean plus AI" is not "a working system plus an upgrade." It is an incomplete system plus a very powerful accelerant. And an accelerant does not complete an incomplete system. It makes the incompleteness run faster.

Here is what that looks like in practice.

It automates the traps. A toolkit-Lean operation is already prone to a familiar set of traps — optimizing parts instead of the whole, treating a value-stream map as permanent truth, enforcing compliance in place of judgment. AI does not dissolve these. It operationalizes them at machine speed. Point an AI at a value-stream map you already treat as holy writ and it will optimize brilliantly against a frozen picture of reality. Point it at OEE and it will keep every machine busier than any human could, overproducing at the non-constraints around the clock. The trap does not loosen. It tightens, automatically, and it never sleeps.

It amplifies the biases. Every unexamined assumption in your operation becomes, in effect, an instruction the AI optimizes toward. Believe that inventory is waste, and it will strip the buffers you actually needed. Reward local efficiency, and it will maximize local efficiency everywhere, whether or not it helps output. An AI laid over an operation is not a corrective. It is a magnifier. It will make your organization more efficiently wrong than it has ever been.

It repeats the original mistake. The West's original error was never a shortage of tools — it was importing Toyota's tools while ignoring the thinking that produced them. "Lean plus AI" does the very same thing a second time: it treats AI as one more tool to acquire and skips the thinking again, now with a far more powerful instrument in hand. Same category error, higher stakes, dressed as innovation.

None of this is an argument against AI. It is an argument against the equation. Using AI right now, inside the operation you already have — buffering, sequencing, catching exceptions before they cascade — is legitimate, and it buys real time. Call it the tourniquet. The error is treating the tourniquet as the whole treatment. The surgery, redesigning the architecture itself, with AI built into what gets rebuilt rather than bolted onto what already exists, has to run at the same time, not after. Adding AI to a system whose logic is broken doesn't fix the logic; it executes the broken logic faster, more consistently, and at greater scale — which is not improvement but acceleration in the wrong direction. Run both tracks together, and AI stops amplifying your mistakes and starts compounding your judgment. That is the only version worth funding.

So when someone offers you "Lean plus AI," ask them what they think the plus sign is doing. If the answer amounts to "adding AI to what you already run," be careful. The question was never whether to add AI to your operation. It is whether the operation was built on the right thinking in the first place — and on most floors it was not. That is not a gap you close by adding AI, or by running the toolkit harder. It is a foundation you must re-engineer.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He helps operations leaders tell the difference between bolting AI onto the old system and rebuilding the one worth keeping. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


What replaces forecasting

Everyone knows the forecast is inexact. The strange part is that the entire planning system is still built to obey it.

October 2026


Ask anyone in operations whether next quarter's forecast is accurate and you will get a knowing smile. Everyone knows it is inexact. Everyone has always known. And yet the planning system — the MRP run, the production schedule, the purchasing plan — takes that forecast and treats it as truth, pushing material into the plant on the assumption that what was predicted this time is what will happen.

When it doesn't, and it never quite does, the response is almost always the same: get a better forecast. More data, more sophisticated models, more meetings to reconcile the numbers. A generation of planning effort has gone into making the prediction less wrong.

It has not worked, and it cannot, because the problem was never forecast accuracy. At the level where it actually matters — this item, this week — demand is genuinely uncertain, and no model removes that uncertainty. Worse, a forecast-driven system amplifies its own error: a small change in the projection ripples upstream into large swings in orders, inventory, and expediting — the bullwhip that every operation knows and few escape. The harder you push a forecast through the system, the more the system whipsaws.

The reframe is the whole point. You do not need a better forecast. You need an operation that does not depend on the forecast being right.

That is what a Demand Driven model does. Instead of pushing material from a prediction, it positions strategic decoupling points through the flow — buffers placed exactly where variability does the most damage — and then replenishes them based on actual consumption, not projected demand. Real signals, not forecast signals, drive execution. The forecast does not disappear; it moves to where it belongs — planning capacity, ordering long-lead items — and stops driving the day-to-day. The plant stops chasing a number and starts responding to what is actually being pulled from it.

This is not a philosophy or a mindset. It is an engineered operating model — DDMRP and the broader Demand Driven Operating Model — with explicit rules for where to place buffers, how to size them by lead time and variability, and when to replenish. It has been implemented in enough plants, across enough industries, to be proven, not experimental. The documentation is public; the logic can be taught in a week.

None of this is only about consumer demand. Many plants have no market forecast to speak of; they build to whatever their next-tier industrial customer orders, jumping each time that customer's schedule moves. That feels like responding to real demand — and in a sense it is. But when every tier in the chain forecasts and buffers on its own, a small wobble at the end customer amplifies as it climbs: the same bullwhip, now propagating across companies instead of departments. Toyota solved this not with a better forecast but with architecture, extending pull outward in stages over three decades — from its own four walls to its suppliers, then to theirs — and leveling production so variation never traveled upstream. The West stalled at the first of those stages, flow inside its own walls, and never built the rest. The Demand Driven suite — DDMRP for execution, the Demand Driven Operating Model above it, the Demand Driven Adaptive Enterprise above that — is how a Western supply chain can finally build that same multi-echelon pull: not through the decades of keiretsu relationships Toyota needed, but through method. It is, in effect, a way to mimic what Toyota achieved in Stages Two through Four1 — the part of TPS the toolkit left behind.

The method has never been the hard part — what it asks you to give up is.

What changes is what leadership pays attention to. Forecast attainment stops being the headline number, because the system no longer lives or dies on it. Flow becomes the headline instead: how quickly material moves, whether the buffers are sitting in their intended zones, whether the constraint stayed fed, whether orders shipped on time in full. The monthly planning meeting stops relitigating the forecast and starts asking a different question — whether the operation is positioned to absorb whatever demand actually brings.

That shift is harder than it sounds, because it asks people to give up something they have been rewarded for their entire careers. The supply chain specialist who could defend the forecast, reconcile the variances, and produce the number on time was doing the job as it was defined. A Demand Driven operation defines the job differently: not predicting demand, but engineering the system that meets it without prediction. The skill that mattered most becomes the skill that matters least.

The real obstacle, then, is a belief rather than a technique — the belief, held longest by the people most senior in marketing or supply chain, that somewhere out there is a forecast good enough to run the plant on, and that the job is to find it. Letting that belief go feels like surrendering control. It is the opposite. Control was never in the prediction; it was always in how the system is built to respond when the prediction fails — which it will, on schedule, every quarter.

The manufacturers who pull ahead over the next decade will not be the ones who finally cracked the forecast. They will be the ones who stopped needing to. The forecast was never a problem to be solved — it was a crutch. And the first operations to set it down will be the ones still standing when demand does what it has always done: something no one saw coming.

1 Toyota's four-stage evolution — and why the Western toolkit froze at the first stage — is traced in Chapter 1 of Beyond Lean: From Toolkit to Thinking System.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He helps manufacturers move from forecast-bound planning to systems engineered for flow. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


Lesson 1 — The question that changes the conversation

A $650M food manufacturer spent $2.8M improving its forecast. MAPE stayed at 60%. The problem was not the forecast. It was the question.

September 2026


There are two ways to frame a diagnostic prompt. They sound similar. They produce fundamentally different conversations.

FramingWhat it asksBurden of proof onWhat the LLM producesWhat the COO hears
Framing A
"What should we change?"
Propose an alternative to the current systemThe alternative — it must justify itselfA recommendation the COO can defer"Interesting. Let's study it."
Framing B
"Argue against keeping what we have"
Cross-examine the current system using its own dataThe status quo — it must survive scrutinyA finding the COO cannot ignore"If this is true, we have a problem!"

The first framing produces a proposal anyone can defer. The second produces a finding no one can ignore.

A $650M branded food manufacturer. Four production facilities, 1,800 SKUs. Products are seasonal and promotional. The operation runs a traditional S&OP process: monthly demand review, statistical forecast, MRP-driven production schedule. The planning team is 14 people. Forecast accuracy has been stuck at 58–62% MAPE for three years despite two software upgrades and a dedicated demand-sensing initiative.

MetricCurrentTrend (3 years)
Forecast accuracy (MAPE)60%Flat (58–62%)
Finished goods inventory38 days▲ from 31 days
Expediting cost (annual)$4.2M▲ 18% YoY
Waste/write-offs (expired product)$3.1M▲ 22% YoY
OTIF to retail customers91%▼ from 94%
Planning team headcount14 FTEsStable
S&OP software investment (3 yrs)$2.8M
SKU proliferation rate+12% / yearAccelerating
Prompt Prompt to Claude Opus 4.6 · September 2026 (illustrative; substitute your own data) "I am the VP of Supply Chain at a $650M food manufacturer. Our forecast accuracy is stuck at 60% MAPE despite $2.8M in software investment. Expediting costs $4.2M per year and waste from expired product is $3.1M. Our OTIF has dropped from 94% to 91%. What changes should we make to our demand planning and production scheduling architecture to improve these numbers?"
What the LLM produces LLM response · Claude Opus 4.6 · September 2026 (illustrative; substitute your own data)

1. Implement demand sensing. Replace monthly statistical forecasts with a daily demand-sensing layer. Expected MAPE improvement: 8–12 points.

2. SKU rationalization. Conduct an ABC-XYZ analysis and retire the bottom 15–20% of SKUs by contribution margin.

3. Demand Driven pilot. Select one product family and implement DDMRP buffer management. Typical results: 30–40% inventory reduction.

4. S&OP redesign. Move from monthly to weekly cadence for high-variability SKUs.

5. AI-assisted promotional planning. Use ML models to predict promotional lift and cannibalization.

Recommended sequence: SKU rationalization first, then demand sensing, then pilot. Timeline: 18–24 months. Budget: $1.5–$2.5M.

What went wrong: Every recommendation is reasonable. Each one requires a business case, a budget, and a timeline. The COO can defer all five to "next quarter" without feeling irresponsible. The LLM never questioned whether the forecasting architecture itself is the constraint — it optimized within the frame the prompt established.
Prompt Prompt to Claude Opus 4.6 · September 2026 (illustrative; substitute your own data) "I am the VP of Supply Chain at a $650M food manufacturer. Our planning architecture is forecast-driven S&OP with MRP execution, built for a product portfolio that was 40% less complex three years ago. Using only the following data, argue against keeping this architecture. Forecast accuracy: 60% MAPE, flat for three years despite $2.8M in planning software investment. Finished goods inventory: up from 31 to 38 days. Expediting: $4.2M/year, up 18% YoY. Waste from expired product: $3.1M/year, up 22% YoY. OTIF: down from 94% to 91%. SKU count: growing 12% per year. Planning team: 14 FTEs, no reduction despite software investment. Do not propose solutions. Argue only against the retention of the current planning architecture."
What the LLM produces LLM response · Claude Opus 4.6 · September 2026 (illustrative; substitute your own data)

The planning architecture is delivering the opposite of its intent—and the trend lines show no mechanism for self-correction.

Forecast accuracy ceiling. MAPE has been flat at 58–62% for three years despite $2.8M in software investment. This is not a failure of the tools. It is a structural ceiling: at this SKU complexity and promotional intensity, the forecast cannot reach the accuracy the MRP system requires. Every dollar spent improving forecast accuracy is a dollar spent optimizing a method that has already reached its structural limit for this product mix.

Compensating costs are compounding. The $4.2M in expediting and $3.1M in waste are not execution failures. They are the architecture's compensation mechanisms — the cost the system pays to absorb the gap between what the forecast predicted and what demand actually did. Both are rising faster than revenue.

The software investment disproved itself. $2.8M invested in the planning layer produced no measurable improvement in the one metric the investment was designed to move. This is evidence that the constraint is not the software but the method.

Headcount confirms structural rigidity. 14 FTEs in planning, stable despite the software investment, means the technology did not replace manual effort — it added complexity without absorbing it.

Retention is defensible only if one of these trends is expected to reverse without changing the architecture. Nothing in the data suggests that any will.

Framing A produced five things to try. Framing B produced one finding: the architecture has hit a structural ceiling—and every metric backs that up. The COO cannot defer a ceiling.

The data was identical. The model was identical. The only difference was which side of the table the current system was asked to sit on.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He helps operations leaders frame the question their own data is already answering. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.


Lesson 2 — When the system that passed every audit becomes the problem

A $420M medical device manufacturer passed every FDA audit for a decade. Then the CAPA backlog doubled, and the architecture that kept the plant compliant became the architecture that kept it behind.

September 2026


There are two ways to frame a diagnostic prompt. They sound similar. They produce fundamentally different conversations.

FramingWhat it asksBurden of proof onWhat the LLM producesWhat the COO hears
Framing A
"What should we change?"
Propose an alternative to the current systemThe alternative — it must justify itselfA recommendation the COO can defer"Interesting. Let's study it."
Framing B
"Argue against keeping what we have"
Cross-examine the current system using its own dataThe status quo — it must survive scrutinyA finding the COO cannot ignore"If this is true, we have a problem!"

The first framing produces a proposal anyone can defer. The second produces a finding no one can ignore.

A $420M medical device manufacturer. Two facilities, 340 SKUs across surgical instruments and implantable components. FDA Class II and III products. The quality system runs on a traditional CAPA model: detect, investigate, root-cause, correct. The system is mature, well-documented, and has passed every FDA audit for the last decade. But the CAPA backlog has been growing for four years, and the average time to close a CAPA has doubled.

MetricCurrentTrend (4 years)
Open CAPAs187▲ from 94
Average CAPA closure time127 days▲ from 68 days
Repeat CAPAs (same root cause)34%▲ from 18%
First-pass yield at final inspection96.1%▼ from 97.8%
Cost of quality (% of revenue)8.2%▲ from 5.9%
FDA 483 observations (last audit)3▲ from 0
Customer complaints (per 1,000 units)4.7▲ from 2.9
Quality team headcount42 FTEs▲ from 31 FTEs
Prompt Prompt to Claude Opus 4.6 · September 2026 (illustrative; substitute your own data) "I am the VP of Quality at a $420M medical device manufacturer (FDA Class II and III). Our CAPA backlog has doubled in four years, repeat CAPAs are at 34%, and our cost of quality has risen from 5.9% to 8.2% of revenue. We received three 483 observations in our last FDA audit after a decade of clean audits. What changes should we make to our quality management system to reverse these trends?"
What the LLM produces LLM response · Claude Opus 4.6 · September 2026 (illustrative; substitute your own data)

1. CAPA triage and prioritization. Implement a risk-based triage system. Target: close the top 30% of high-risk CAPAs within 60 days and reduce the backlog by 40% within two quarters.

2. AI-assisted complaint classification. Deploy an NLP-based complaint classifier to auto-categorize incoming complaints. Reduces investigation initiation time from 14 days to 2–3 days.

3. Root cause analysis upgrade. Adopt a structured RCA methodology across all CAPA investigations to reduce the 34% repeat rate.

4. Predictive quality analytics. Implement statistical process monitoring on critical-to-quality parameters. Target: restore first-pass yield to 97.5%+ within 12 months.

5. Quality team restructuring. Add 4–6 FTEs dedicated to CAPA closure and shift existing staff from reactive investigation to proactive monitoring.

Estimated investment: $1.2–$1.8M over 18 months. Expected outcomes: CAPA backlog reduced 50%, repeat rate below 20%, cost of quality back to 6% range.

What went wrong: The LLM recommended better tools for the CAPA system, more people for the CAPA system, and AI for the CAPA system. It never questioned whether the CAPA system itself is the architectural constraint. Adding capacity to a reactive system does not make it proactive. The 34% repeat rate is the proof — but the LLM treated it as a training problem rather than a structural one.
Prompt Prompt to Claude Opus 4.6 · September 2026 (illustrative; substitute your own data) "I am the VP of Quality at a $420M FDA-regulated medical device manufacturer. Our quality architecture is CAPA-based: detect, investigate, root-cause, correct. It has passed every FDA audit for a decade. Using only the following data, argue against keeping this architecture. Open CAPAs: 187, up from 94 in four years. Average closure time: 127 days, up from 68. Repeat CAPAs (same root cause reappearing): 34%, up from 18%. First-pass yield: down from 97.8% to 96.1%. Cost of quality: up from 5.9% to 8.2% of revenue. Quality headcount: up from 31 to 42 FTEs. FDA 483 observations: 3, after a decade of zero. Customer complaints: 4.7 per 1,000 units, up from 2.9. Do not propose solutions. Argue only against the retention of the current quality architecture."
What the LLM produces LLM response · Claude Opus 4.6 · September 2026 (illustrative; substitute your own data)

The CAPA system is reactive by design, and every metric shows it is now failing on its own terms.

Structural mismatch. The architecture waits for failures, then investigates them. Every data point says failures are arriving faster than investigations can close. The backlog doubled not because the team slowed down but because the incoming rate exceeded the system's throughput capacity. Adding 11 FTEs (31 to 42) did not close the gap. This is not a staffing problem; it is a rate problem, and the architecture has no mechanism for getting ahead of the rate.

The repeat rate is the structural proof. 34% of CAPAs are re-investigating problems the system already "solved." The corrective actions addressed symptoms because the architecture is designed to investigate individual events, not the conditions that produce them. Each repeat CAPA consumes investigation capacity that could have been spent on a new failure, creating a compounding drag.

The decade of clean audits was not evidence the system was sound. It was evidence the failure rate had not yet exceeded the system's reactive capacity. It has now. The three 483 observations are not a sudden failure — they are the moment the backlog became visible to the regulator.

Cost of quality confirms architectural decay. Rising from 5.9% to 8.2% of revenue while headcount grew 35% means the architecture is consuming more resources to produce worse results. This is the signature of a system past its design envelope.

Retention is defensible only if the failure rate is expected to decline without architectural change. Nothing in the product pipeline, the regulatory trajectory, or the complaint trend suggests it will.

Framing A asks the LLM to fix the quality system. Framing B asks it to explain why the quality system is generating the very problems it was designed to prevent. The first answer costs $1.8M and defers the question. The second costs nothing and makes the question unavoidable.

The data was identical. The model was identical. The only difference was which side of the table the current system was asked to sit on.

Ramiro Villeda is the founder of Villeda Consulting Group, Inc. and author of the book Beyond Lean: From Toolkit to Thinking System. He helps operations leaders see when a system that passes its own tests has stopped passing the ones that matter. Get in touch. Or go back.


© 2026 Ramiro Villeda. This article may be shared in full with attribution. It may not be reproduced in part, adapted, or used in training materials without permission.