T-shirt sizing is a relative estimation technique that uses ordinal labels like XS through XL to describe the effort, scope, and uncertainty of work instead of hours or story points. Teams use it for early discovery, roadmap and release planning, and triaging large backlogs before deeper refinement. It gives direction, not dates: for actual delivery forecasts, pair it with cycle time or throughput data rather than converting labels straight into a calendar.
TL;DR:
- T-shirt sizing is best calibrated with completed work to ensure relative effort and complexity are accurately reflected for future items.
- Use t-shirt sizes during high-level backlog prioritization and sequencing, not for sprint planning or committed delivery dates.
- Incorporate a short rationale for each size and treat XL as a signal to split or spike an item to manage uncertainty effectively.
- Base delivery forecasts on cycle time or throughput data, not on size labels, and report ranges or percentiles for more realistic estimates.
- Maintain a consistent scale anchored to reference examples, and revisit sizing calibration regularly as team understanding and work evolve.
Table of Contents
- What is t-shirt sizing and how do teams set a scale?
- When should you use t-shirt sizing instead of story points?
- How do you run a t-shirt sizing session with your team?
- How do you turn sizing sessions into a repeatable process?
- What benefits and traps come with t-shirt sizing?
- What does this look like on a real backlog?
- How do you convert sizes into honest forecasts?
- Where a phase-0 planning tool fits into this workflow
- The rules I actually follow when sizing
- Make calibration and rationale capture part of your process
- Sources
- FAQ
What is t-shirt sizing and how do teams set a scale?
T-shirt sizes work because they sidestep the false precision of hours. Labels like XS, S, M, L, and XL (some teams add XXS or XXL for extreme outliers) describe an item's effort, complexity, and uncertainty relative to other items, not a fixed duration. Nobody argues that an L "means" 13 hours, which is exactly the point. The scale forces a conversation about scope instead of a guess dressed up as math.
The technique only works when it's calibrated to real, completed work. Before sizing anything new, pick two or three finished items your team already agrees on and label them S, M, and L. These become your reference stories: when someone proposes a size for a new item, the question isn't "how big does this feel," it's "is this bigger or smaller than the M we shipped last quarter." According to Easy Agile's guidance on t-shirt sizing, anchoring to completed reference work and recording the reasoning behind each size choice are what separate useful sizing from guessing.

Each label typically reflects three things at once: how much work is involved, how complicated the solution is, and how much you don't yet know. An M with high uncertainty and an M with high complexity might feel like the same size on the surface, but they carry different risks, which is why capturing a short rationale matters more than the label itself.
Common misuses to avoid:
- Treating a size as a promise of hours or days instead of a relative comparison.
- Reusing another team's scale without recalibrating to your own reference work.
- Skipping the rationale and leaving only a letter on the ticket.
- Letting one persuasive voice set the size instead of using independent estimates.
When should you use t-shirt sizing instead of story points?
T-shirt sizing and story points solve different problems, and the mistake most teams make is picking one and using it everywhere. According to Asana's overview of t-shirt sizing, t-shirt sizing suits high-level comparison and prioritization during roadmap or release planning, while story points are built for the more refined estimates a team needs before committing to a sprint.
Use t-shirt sizing when:
- You're scanning a large, unrefined backlog and need a fast sense of relative scope.
- You're in program or release planning and need to sequence epics before anyone has written acceptance criteria.
- Stakeholders need a rough comparison between initiatives, not a delivery commitment.
Use story points when:
- The team is planning a sprint and needs enough shared understanding to commit.
- The work has been broken down into stories small enough to discuss in detail.
- You need a number that feeds into velocity tracking for that team specifically.
Some teams bridge the two with a rough conversion table (XS to a few points, XL to many more), which is fine as a provisional bridge but dangerous as a permanent habit. Mark any such conversion "provisional" and replace it with a real point estimate once the team actually refines the item before sprint commitment. Never compare t-shirt sizes across teams: an L on one team's backlog has no relationship to an L on another's, because the reference work behind each scale is different.
Pro Tip: If you catch yourself doing arithmetic on t-shirt sizes, averaging them or adding them up, stop. That math hides exactly the qualitative differences the labels were meant to preserve.
How do you run a t-shirt sizing session with your team?
Two facilitation patterns cover almost every situation: a small-set session for a manageable batch of items, and affinity mapping with buckets when the backlog is too large to discuss one by one.
Small set pattern (20 to 30 minutes, 10 to 20 items):
- Choose 10 to 20 items that need sizing and pull up your reference stories for S, M, and L.
- Have everyone size the first item independently and reveal simultaneously (cards, a Jira field, or a quick poll all work) so no one anchors on the first voice spoken aloud.
- If the group agrees or is one size apart, record it and move on.
- If sizes diverge widely, let the highest and lowest estimator explain their reasoning in one sentence each, then re-vote once.
- Capture the final size plus a short rationale directly on the ticket before moving to the next item.
Affinity plus buckets pattern (large backlogs):
- Lay out all items (physical cards or a digital board) and ask the group to arrange them from smallest to largest by feel, with no discussion yet.
- Draw soft boundaries across the line to mark XS through XL bands.
- Allow one pass where anyone can challenge an item's placement, but only if they can state what specific information would change the estimate.
- Move challenged items based on that one pass, then close the session.
This pattern comes from Parabol's t-shirt scale template, which frames the one-pass challenge rule as the mechanism that keeps affinity sizing fast instead of turning into a debate about every card.
A few things make either pattern work better:
- Simultaneous reveal beats a round-robin, since the first number spoken tends to anchor everyone after it.
- The people doing the work should size the work, not a manager sizing on their behalf.
- Tooling can stay simple: a Jira custom field for size, physical index cards, or a shared spreadsheet all do the job.
Assign two roles before you start: a facilitator who keeps the group moving and enforces the one-pass rule, and a recorder who captures the size and the one-sentence rationale so nobody has to reconstruct the reasoning later.
How do you turn sizing sessions into a repeatable process?
A sizing session only pays off if the outputs survive past the meeting. Four habits keep the practice consistent over time.
Start by defining your team's scale explicitly and writing it down somewhere visible, then anchor it to two or three items you've actually completed. Revisit that anchor set whenever the team's context shifts, such as picking up a new type of work or bringing on new members who weren't part of the original calibration.
Second, record more than the letter. Every sized ticket should carry a one-sentence rationale covering the thing that drove the size: an unclear dependency, a known technical risk, or a gap in requirements. This turns a static label into a note future refinement can actually use.
Third, treat XL as a decision point, not a valid backlog item. A discussion thread on Scrum makes the case directly: teams should use feature-level sizes to compare scope, then refine a subset further before setting release targets, rather than letting an oversized item sit in the backlog labeled XL indefinitely. If an item earns an XL, split it or spin up a discovery spike to shrink the uncertainty before it goes anywhere near a sprint.
Finally, build sizing into your existing rituals instead of treating it as a special event. Size new items during backlog grooming, revisit epics during roadmap sessions, and schedule a periodic review, quarterly is reasonable for most teams, where you compare the sizes you assigned against what actually shipped.
- Write the scale down and anchor it to completed work everyone recognizes.
- Capture a one-sentence rationale on every sized ticket.
- Split or spike any item that lands on XL.
- Review past sizes against real delivery data on a set schedule.
Pro Tip: Keep your anchor examples in the same place your team already looks, a pinned document, a wiki page, or a tool built for exactly this, so recalibration takes minutes instead of a meeting.
What benefits and traps come with t-shirt sizing?
Done well, t-shirt sizing is fast and inclusive: non-technical stakeholders can participate in a conversation about scope without needing to understand story point math, and a room full of people can size a quarter's worth of epics in under an hour. It also surfaces disagreement early, before anyone has invested real time in a plan, which is exactly when disagreement is cheapest to resolve.
The traps are predictable and almost always social rather than technical.
- Stakeholders hear "M" and ask for a delivery date, treating a rough label as a schedule commitment.
- Someone converts sizes to dates by multiplying a made-up number of days per letter.
- Teams compare their scale to another team's as though the letters meant the same thing everywhere.
Each has a direct fix. When a stakeholder pushes for a date, Scrum.org's forum discussion points to explaining the difference between an estimate and a forecast, then using refined analysis and historical delivery data to answer with a range instead of a point. Mark any size-to-points conversion as provisional. And require a short rationale on every size so the "what would change this" question has an answer ready when someone asks.
What does this look like on a real backlog?
- A retail team sizing a checkout revamp splits the epic into "update payment provider" (M), "redesign cart page" (S), "add saved payment methods" (L), and "support split payments" (XL). The XL is flagged immediately rather than dropped into a sprint.
- That XL, split payments, becomes a two-week discovery spike. The spike surfaces a hard dependency on the payment provider's API limits, and the team resizes the original item down to L once the unknowns are resolved.
- Across program increment planning with three teams, sizing lets planners sequence "redesign cart page" before "split payments" without ever pretending an L on the checkout team equals an L on the fulfillment team; sequencing runs on relative priority within each team's own backlog, not on cross-team point totals.
How do you convert sizes into honest forecasts?
A t-shirt size answers "how big is this relative to other things," not "when will it ship," and the two questions need different data. According to the Agile Alliance's guidance on explaining dates from historical data, the right inputs for a delivery forecast are cycle time and throughput, the actual pace at which your team finishes work, not a label converted through an invented ratio.
A commonly cited approach is Monte Carlo forecasting built on cycle time or throughput. Collect 8 to 12 weeks of weekly throughput data, run many simulations of how that pace could play out going forward, and report the results as percentiles, for example "85% of simulated futures finish by week N," rather than a single date, a method described in the Agile Alliance's forecasting article.
If you need a rough size-to-points bridge for prioritization purposes, keep it explicitly provisional and replace it with a real point estimate once the team refines the item before sprint planning.
- Never promise a date from a t-shirt label alone.
- Use 8 to 12 weeks of throughput as your forecasting baseline.
- Report ranges and percentiles, not single deterministic dates.
- Treat any size-to-points conversion as temporary scaffolding, not a final number.
Where a phase-0 planning tool fits into this workflow
Sizing works best when reference anchors and rationale notes are easy to find later, which is exactly the gap a structured planning tool can close before you ever open a sprint board. A phase-0 planning tool builds a connected model of a business idea (scope, market, features, requirements) starting from a one-line description, and that same structure gives founders a place to store completed-reference examples and the one-sentence rationale behind early scope decisions, so early sizing conversations aren't lost once the backlog grows. For a team still defining its product before a formal backlog exists, that kind of phase-0 structuring can matter more than an estimation ritual: sizing has nothing to calibrate against until the shape of the work is clear. Once you're managing an active backlog inside Jira or a similar tool, running sizing sessions there directly still makes sense. Exports to Notion and Confluence on advanced plans let that early structuring feed straight into whatever project tooling the team already uses.
The rules I actually follow when sizing
Three defaults save most arguments before they start: keep the scale to five labels, anchor every size to real completed work, and treat any XL as a signal to split the item or spike it, never as a backlog entry you plan around. If the room is split after one re-vote, record the higher size provisionally and move on. Debating a letter for ten minutes costs more than being wrong by one size.
— Karl
Make calibration and rationale capture part of your process
Klaritea turns a one-line idea into a structured model covering scope, market, features, and requirements, giving teams a place to keep the reference examples and rationale notes that make sizing sessions repeatable instead of reinvented each time. 
If your team is still shaping what the backlog will even contain, it's worth structuring that thinking before your first sizing session rather than after. You can start on the Free plan or view the Klaritea and Pro plans to see which fits a team preparing for its next release.
Sources
- Agile estimation techniques: A deep dive into t‑shirt sizing | Easy Agile
- T Shirt Sizing: Agile Estimation & Capacity Planning | Asana
- Explain dates to anyone with forecasts based on your historical data | Agile Alliance
- T-shirt scale template | Parabol
- Scrum
For teams working through how sizing connects to sprint-ready stories, this piece on turning requirements into sprint-ready stories walks through when to size at a high level and when to refine further, and this guide to opportunity assessment covers facilitation steps useful for picking reference work. Teams that need to split an XL into something testable can also look at user story mapping for an MVP for a practical slicing approach, and teams scaling facilitation practices across more teams may find the engineering and process support at Red Rock Webscapes useful.
FAQ
How is t-shirt sizing used in Agile story points?
T-shirt sizing and story points are separate techniques, though some teams use a provisional conversion table to bridge them during early planning. The safer practice is to use t-shirt sizes for roadmap and release-level comparisons, then assign real story points once the team refines the item ahead of sprint commitment.
What does t-shirt size mean in Jira?
In Jira, a t-shirt size is usually a custom field (XS through XL) attached to an epic or story to show relative effort or scope at a glance. It's a labeling convention your team sets up rather than a built-in Jira estimation unit, so the scale and its meaning are defined locally by each team.
What software can I use to estimate t-shirt sizes in Agile projects?
Many teams size directly in the tool they already use for backlog management, adding a custom field in Jira or a similar platform, while others use dedicated estimation templates like the one from Parabol for affinity-based sessions. A phase-0 planning tool such as Klaritea can also hold reference examples and rationale notes earlier, before a formal backlog exists.
How can I estimate t-shirt sizes in Jira?
Add a custom field for t-shirt size to your issue type, then run a sizing session using either the small-set pattern for a short list of items or affinity grouping for a large backlog. Record the size and a short rationale directly in the ticket so the reasoning is visible the next time the item comes up for refinement.
Should stakeholders expect delivery dates from t-shirt sizes?
No, a t-shirt size describes relative scope and uncertainty, not a timeline. For an actual delivery estimate, teams should use cycle time or throughput data to produce a percentile-based forecast instead of converting a label directly into a date.
