Agile Alliance members get free access to the Instant Insights tool. Go to the Comparative Agility service provider page, make sure you’re logged in, and you’ll see a link on the left-hand side.

The following is an AI summary of the event.
The presentation deck is available at the bottom of the page.
Overview
This webinar, Making Continuous Improvement Measurable in Agile, was part of the Comparative Agility series. Hosted by Jorgen Hesselberg and presented by Mooly Beeri, CEO and co-founder of BetterSoftware.dev, the session focused on a practical problem many Agile organizations face: teams run iterations, hold retrospectives, and discuss improvements, but often lack a structured way to measure whether they are actually getting better.
Beeri argued that continuous improvement is central to Agile, but in many frameworks it is only implied rather than explicitly managed. As a result, teams may keep delivering in short cycles without making measurable progress in quality, time to market, engineering effectiveness, or customer outcomes.
Why Agile Improvement Often Stalls
Beeri opened by noting that Agile is widely used and viewed as important, but many organizations still report low proficiency. His explanation was direct: Agile teams often create room for improvement, but they do not always require improvement.
In many teams, retrospectives identify problems but do not produce specific, measurable process changes. Improvement work becomes scattered across many competing ideas: adopting SAFe, changing architecture, increasing unit test coverage, adding AI coding tools, improving CI/CD, introducing DORA metrics, or replacing development tools. Without a shared measurement system, leaders struggle to decide which investments matter most.
The result is what Beeri described as a “spray improvement” pattern: a little budget goes to many initiatives, with no clear evidence that the overall organization is becoming more effective.
The Need to Measure Engineering Effectiveness
The main point of the session was that organizations need a way to measure engineering effectiveness before they can improve it. Beeri presented a model derived from software development lifecycle practices, covering 35 competencies across areas such as analysis, design, implementation, testing, release, and maintenance.
Each competency is scored on a 10-level effectiveness scale. The lower levels represent little or no activity. Middle levels show that a team is performing and measuring a practice. Higher levels show that the team controls trends, fixes issues at the point of injection, experiments with better methods, and influences broader organizational practice.
This gives teams and leaders a shared view of where capability is strong, weak, or uneven.
The Wheel Heatmap
Beeri showed a “wheel heatmap” as the central visualization in the model. Each section of the wheel represents an engineering competency, and the colors show relative strength or weakness. The model also produces an overall effectiveness score out of 100.
The value of the wheel is that it makes improvement priorities visible. Instead of relying on opinion or whoever argues most strongly in a retrospective, the team can see which areas are holding them back. For example, one sample team appeared strong in requirements and architecture, but weaker in implementation areas such as code maintainability, reuse, and third-party code integration.
The model also weights competencies differently, so improving one area may have more impact on overall effectiveness than improving another.
From Retrospectives to Measurable Improvement Plans
A major recommendation was to turn retrospectives into planning sessions for measurable process improvement. Beeri suggested that every iteration should include at least one measurable process enhancement, not just delivery work.
In this model, a team might commit to moving continuous integration from one effectiveness level to the next, improving code maintainability by two levels, or strengthening requirements practices. The improvement is then tracked, implemented, and reviewed in the next cycle.
This changes the retrospective from a discussion about what went wrong into a mechanism for deciding what the team will improve next and how progress will be verified.
Team Ownership With Organizational Direction
Beeri emphasized that teams should design their own improvement paths, but not in an uncontrolled way. Leadership can set a target, such as requiring every team to improve its overall effectiveness score by a certain number of points per quarter or year. Within that target, each team chooses which areas of the model to improve based on its own context.
This balances local ownership with organizational consistency. Teams are not all forced to adopt the same solution, but they are all working within the same measurement system and contributing to measurable improvement.
AI and Engineering Effectiveness
The session also addressed AI in software development. Beeri pushed back on the idea that AI makes the software development lifecycle irrelevant. Instead, he argued that AI makes measurement more necessary.
Organizations are investing in AI tools for requirements, coding, testing, and analysis, but often cannot prove whether those tools are improving outcomes. Beeri recommended using an effectiveness model to decide where AI is likely to help. For example, if a team already has strong requirements practices, replacing requirements work with an AI agent may add little value. If the team struggles with code maintainability, AI support may be more useful.
He also described a cautious approach: benchmark AI against human performance in a specific lifecycle activity, then expand AI usage only when it produces better effectiveness than the current human-led process.
Case Studies
Beeri shared several examples from large software organizations.
At Philips, the model was used across roughly 5,000 software engineers between 2016 and 2018. Beeri reported a 10% to 12% year-over-year reduction in defect volume, nearly 30% overall reduction in defect volume, and about 60% reduction in post-release issues.
At Aptiv, the approach was applied across about 3,500 software engineers starting in late 2020. Beeri reported a 14% to 16% year-over-year reduction in defect count, about 50% overall reduction in defect volume, close to 40,000 late-stage defects eliminated, and a 50% reduction in post-release issues.
At Mercedes, the work focused on a smaller team responsible for a critical component. Beeri said the team achieved about a 90% reduction in defect volume. He clarified during Q&A that this was an unusually strong case, helped by the fact that the team already had useful engineering assets, such as automation, but was not using them in the right sequence or as proper gating controls.
Discussion From the Q&A
During the Q&A, attendees asked how the model is measured. Beeri explained that the lightweight version available through Comparative Agility is a self-assessment. The full version combines questionnaires, deep dives, and coaching support. In that version, a coach works with the team continuously, similar to an ongoing improvement specialist rather than a once-a-year auditor.
Another attendee asked whether too many changes at once could overwhelm teams. Beeri responded that too many simultaneous changes usually signal a lack of prioritization. In his approach, teams focus on only a few improvement initiatives per iteration, chosen because they are expected to move effectiveness fastest.
There was also discussion about defect metrics. Beeri explained that the goal is not only to reduce the total number of defects, but also to find and fix defects earlier in the lifecycle. Higher effectiveness levels are associated with fixing issues at the point where they are introduced, which reduces the cost of correction and lowers customer escapes.
One attendee expected more tool-specific metrics for Jira or Azure DevOps. Beeri’s response was that most teams already have too many dashboards and numbers. A useful metric should have clear thresholds: when it turns red, when it turns green, who owns the response, and what action should follow. Without that logic, adding more numbers rarely changes behavior.
Final Takeaways
- Continuous improvement is not automatic just because a team is working in iterations.
- Retrospectives are weak if they produce discussion without measurable process changes.
- Engineering effectiveness can be measured across specific software lifecycle competencies.
- A shared effectiveness model helps teams choose improvement work based on evidence rather than opinion.
- AI investments should be evaluated against actual improvements in effectiveness, not adopted because they are fashionable.
- The strongest operating model is team-owned improvement within a consistent organizational measurement framework.



