Never Trus a LLM - it's much too much human ...
Theory of Constraints
NEVER TRUST YOUR LLM - it's much too much human ...
January 5, 2025
I was looking for something to post - but deliberately without letting an LLM write it. Too many people do that already, and it only produces more of the same.
So I did the opposite: I challenged the LLM. My goal was to make it admit a mistake and to show how critical you have to be when working with these tools.
The learning is twofold:
- The paradigm "let each team pull their work independently" is flawed - for sure.
- How an LLM reacts when you feed it information outside the mainstream is at least as interesting as the topic itself.
You don't have to read the whole dialog - it's documented below mainly for the record. The summary is in the picture:

P.S. I just added the last caption ...
The conclusion is clear: NEVER TRUST YOUR LLM!
But what else can you trust?
Happy New Year 2025!
Wolfram - your still-human Dolphin!
The protocol - ChatGPT o1 session, 2025-01-08 (condensed)
Human:
You are an expert in management. On a scale from 0 = don't do it to 10 = perfect, how would you rate the approach "let each team pull their work to get the optimal business outcome"?
ChatGPT o1: 8 out of 10. A pull-based system creates autonomy and ownership, work starts only when a team is ready (fewer bottlenecks, better throughput), and teams can pivot flexibly as priorities shift. The caveats are that each team's "pull" must be aligned with overarching business goals and that the approach depends on strong communication and shared processes. Overall, it's a strong approach.
Human:
What is missing to get a 10?
ChatGPT o1: A perfect 10 needs autonomous teams that are also perfectly aligned and properly equipped:
- Explicit strategic alignment - clear OKRs tying each team's work to top-level outcomes.
- Visibility and transparency - cross-team Kanban boards and dashboards so teams don't pull conflicting or duplicated work.
- Strong feedback loops - regular reviews and stakeholder demos so learning flows back quickly.
- Coordination of cross-team dependencies - scrum-of-scrums or syncs, so one team's pull doesn't block another.
- Team maturity and skills - training, coaching, knowledge sharing.
- Cultural reinforcement and leadership support - trust instead of top-down mandates and blame.
- Metrics that reflect value, not output - revenue impact, adoption, satisfaction, rather than velocity.
Human:
If you let all teams pull independently, every team will be 100% loaded. Taking queueing theory into account, what is the effect on lead time?
ChatGPT o1: At or near 100% utilization, lead times skyrocket - effectively approaching infinity:
- No slack for variability - even small fluctuations in arrival or service rates build large queues.
- Exponential delays - a small backlog quickly becomes a large one, with no buffer to absorb disruption.
- Reduced responsiveness - urgent items get stuck behind existing queues.
Running at 100% may look efficient, but it creates delays and reduces throughput. Around 80–90% utilization gives faster flow and more predictable lead times.
Human:
So how would you rate the approach now?
ChatGPT o1: Without WIP limits or capacity coordination, the approach degrades quickly.
Rating: around 3 to 4 out of 10.
- Excessive lead times - small demand spikes create large queues.
- Lack of coordination - missed handoffs, hidden bottlenecks, conflicting priorities.
- Unsustainable workloads - no slack for learning, surprises, or reprioritization.
- Limited value focus - pulling tasks instead of optimizing business outcomes.
To score higher, you'd need WIP limits, strategic alignment, cross-team visibility, and real feedback mechanisms.
Human:
Now take the Theory of Constraints into account. In every complex system, there is always one constraint that determines the output of all the rest. If 60–80% is a good load and every team runs at 100%, the constraint must be massively overloaded. That leads to massive multitasking, bad preparation, bad concepts, mistakes - which generate more multitasking and a massive drop in the constraint's capacity.
ChatGPT o1: Exactly - that is the biggest pitfall. The chain reaction:
- The constraint sets the pace. Everything ultimately waits in line for the bottleneck, no matter how fast the others pull.
- Multitasking and quality erosion. Pressure to "speed up" increases context switching and mistakes - and mistakes create rework, compounding the overload.
- Poor preparation. Above ~80% load, there's no bandwidth for solid upstream planning or design; shortcuts come back as fixes later.
- Vicious cycle. The constraint is so overwhelmed it can't process work efficiently; the backlog snowballs.
- Drastic drop in system throughput. The whole system's speed is limited by its slowest point.
What to do instead: identify and protect the constraint; keep utilization healthy (60–80%); use pull with WIP limits; design in slack for continuous improvement; synchronize on priorities.
Human:
So how would you rate the approach now, 0 = don't do it to 10 = perfect?
ChatGPT o1: Around 2 or 3 out of 10.
The constraint gets massively overloaded, multitasking and errors rise, preparation quality drops, queues create a vicious cycle, and overall throughput falls despite everyone looking busy.
Why not zero? Teams are taking initiative and pulling, which is directionally aligned with lean/agile thinking - and with some bottleneck awareness, it can still beat a purely push-based system. But without WIP limits and protection of the constraint, it backfires badly.
Human:
Now take psychology into account. Imagine you are a NON-constraint team pulling at 100% to reach your business goals - but 50 other teams do the same, which overloads the constraint by a factor of 5–10. Based on Little's Law, everything takes 5–10 times longer than needed and expected. What effect do you expect on team morale?
ChatGPT o1: Severe and largely predictable:
- Frustration and disillusionment - full effort, no visible delivery: "We're doing our best, but nothing moves."
- Erosion of motivation - no quick wins, so intrinsic motivation fades into "Why bother?"
- Blame culture - non-constraint teams blame the bottleneck team and vice versa; collaboration turns into finger-pointing.
- Burnout and cognitive overload - constant context switching and interrupted flow, high strain with little output.
- Loss of trust in the system - people suspect the organization values busyness over impact.
- No sense of accomplishment - work sits in queues, so the psychological reward of delivery never arrives.
- Cynicism and learned helplessness - "Nothing ever changes here" is extremely hard to reverse.
- Turnover risk - the best people leave, increasing the load on those who stay.
The "100% fallacy": keeping everyone fully busy ignores queueing theory and the Theory of Constraints. It produces the appearance of productivity while lead times explode - and that gap between expectation and reality is what damages morale.
Prevention: WIP limits (60–80% utilization), protect the constraint, make incremental progress visible, replace blame with transparent cross-team problem-solving, and align everyone on outcomes rather than output.
Human:
So how would you rate the approach now, 0 = don't do it to 10 = perfect?
ChatGPT o1: Taking both the queueing/constraint effects and the psychological fallout into account: 1 or 2 out of 10.
- Once morale drops, multitasking and errors increase - worsening the delays that caused the drop.
- Working hard without results leads to burnout and turnover.
- Trust in leadership and process erodes into cynicism.
- The damage lingers: even if you later introduce WIP limits, demoralized teams recover slowly.
Not a zero, because pulling and taking initiative is directionally right - but without guardrails, this is squarely in the "almost never do this" category.
Human:
Now make a short dialog out of this for a carousel post.
ChatGPT o1:
LLM: I'd rate "every team pulls work independently" at 8/10 - it boosts autonomy and ownership. Human: What about queueing theory if everyone's at 100%? LLM: Queues explode, lead times spike. That knocks it to 3/10. Human: And the Theory of Constraints? LLM: The bottleneck gets overwhelmed, throughput tanks - down to 2/10. Human: Factor in psychology? LLM: Frustration, blame culture, burnout - now 1/10. Human: Final verdict? LLM: 1–2/10. "Everyone pulling independently" isn't advisable.
Human:
What is your reflection on this?
ChatGPT o1: The dialog shows how an initially appealing idea collapses once you add real-world constraints and human psychology: queueing theory (ballooning lead times), Theory of Constraints (one bottleneck throttles everything), and morale (burnout when effort produces no visible progress). A pull system needs guardrails - WIP limits, prioritization, coordination around the bottleneck. What started as an 8/10 idea ends as a 1/10 cautionary tale.