Velocity is not a performance metric. Story points do not tell you if your team is delivering value. Here are the agile metrics that actually correlate with delivery health on web development teams and how to use them to drive improvement.
Most web development teams that practice agile report on velocity every sprint. Velocity is a measure of how many story points a team completes per sprint and it is, with rare exceptions, a poor indicator of anything that matters to the business. It is sensitive to estimation accuracy, which varies widely between teams and shifts within a team over time. It is easy to game by inflating estimates. It does not measure whether what was delivered actually worked for users. It says nothing about quality, lead time, or the rate at which the team is improving. The reason web teams continue to track velocity is inertia and tooling defaults most agile tools surface velocity prominently because it is easy to calculate, not because it is meaningful. The metrics that actually correlate with delivery health on web teams are different, and most of them require slightly more effort to track than story point counts.
Cycle time is the elapsed time from when work starts to when it is delivered in a web context, from when a developer picks up a ticket to when the associated code is deployed to production and confirmed working. It is the single most informative metric for diagnosing delivery system problems. High cycle time is almost always a symptom of one of a small number of root causes: work items that are too large to complete in a single sprint, excessive work in progress that causes context-switching and queue buildup, review and approval bottlenecks that hold completed code waiting for sign-off, or inadequate testing and deployment automation that makes the path to production slow and risky. Because cycle time integrates all of these failure modes into a single number, trending it over time reveals whether the delivery system is improving or deteriorating without requiring agreement on which specific failure mode is the priority.
Deployment frequency how often the team deploys working software to production is one of the most powerful signals of delivery system health for web teams. Teams that deploy frequently (daily or multiple times per day) have shorter feedback loops, lower deployment risk per release, and faster recovery paths when something goes wrong. Teams that deploy infrequently (weekly, monthly, or longer) accumulate deployment risk with each batch of changes and typically discover that the infrequency itself is a root cause of quality problems rather than a protection against them. Change failure rate the percentage of deployments that require rollback, hotfix, or immediate remediation is the quality complement to deployment frequency. A team with high deployment frequency and low change failure rate has a fundamentally healthy delivery system. A team with high deployment frequency and high change failure rate has a continuous integration problem. These two metrics together tell a more accurate story than either one alone.
The DORA (DevOps Research and Assessment) metrics deployment frequency, lead time for changes, change failure rate, and mean time to restore were developed through research across thousands of software teams and represent the most empirically grounded framework for measuring delivery performance available. For web development teams, applying DORA requires adapting the definitions slightly. Lead time for changes, in the web context, is measured from code commit to production deployment not from ticket creation, which introduces planning and prioritization time that is not a property of the delivery system itself. Mean time to restore is the elapsed time from when a production incident is detected to when service is fully restored. Tracking this metric creates organizational pressure to invest in observability, incident response, and deployment rollback capability investments that most web teams defer until they experience a serious incident.
The fundamental problem with velocity as a team performance metric is that it is entirely self-referential. A team can have high velocity and be delivering work that takes months to reach production, has a high defect rate, and creates significant maintenance burden. A team can have low velocity and be delivering clean, fast-deployable software that delights users. Neither velocity number tells you which situation you are in. Story points compound this problem by adding a layer of estimation that creates the appearance of precision without corresponding accuracy. Research on software estimation consistently shows that teams' estimates have wide error bands and that aggregating them into velocity figures does not improve predictive accuracy in a reliable way. The appropriate use of story points is as a planning heuristic a rough guide to what the team can take on in a sprint not as a performance measurement that gets reported upward as evidence of productivity.
Business outcome metrics customer satisfaction, revenue, feature adoption are lagging indicators of delivery quality. They tell you whether what was delivered was valuable, but they do so slowly and with enough confounding factors that it is difficult to connect specific delivery decisions to specific outcomes. Delivery metrics like cycle time, deployment frequency, and change failure rate are closer to leading indicators they predict whether the team is capable of delivering outcomes, and they respond to changes in ways of working much faster than business metrics do. The most useful measurement architecture for web teams combines both: a small set of delivery metrics that the team owns and reviews in retrospectives, and a small set of business outcome metrics that connect delivery to the value the organization cares about. Teams that only track delivery metrics optimize their delivery machinery without knowing if the output is valuable. Teams that only track business outcomes cannot connect cause to effect.
A useful metrics dashboard for a web development team has at most five to seven metrics, each of which has a clear owner and a clear interpretation. Cycle time (median and 85th percentile, trending over time), deployment frequency (deployments per week), change failure rate, and mean time to restore cover the DORA framework. Add sprint goal achievement rate the percentage of sprint goals fully met, not the percentage of stories completed as a team health indicator. Consider adding a customer-facing quality metric like user-reported defect rate or error monitoring event volume. What should not be on the dashboard: raw velocity, story point counts, individual developer productivity metrics (which are both inaccurate and damaging to team culture), and percentage-of-time-in-meetings or any other proxy that teams learn to optimize without improving delivery. The test for any metric is simple: if the team improved this number without improving their actual delivery, would anyone know? If the answer is yes, the metric is gameable and should not be tracked as a performance indicator.
Talk to our experts about how we can help your organization apply these insights in practice.