#Tagly
Measurement

How to actually test whether your hashtags are working

Most hashtag "testing" compares two posts and draws a conclusion. That is not enough data to distinguish a real effect from noise, and it is why so much hashtag advice contradicts itself.

7 min readUpdated 28 August 2026By Tagly

Hashtag advice is unusually contradictory, and the reason is measurement. Individual post performance varies enormously for reasons unrelated to tags — the hook, the topic, the day, the thumbnail, what else was happening in the feed. Comparing one post to another and attributing the difference to hashtags is a coin flip described as a finding.

Testing properly is not complicated, but it does require a bit of discipline about what you measure and how many observations you need.

Measure the right thing

The critical distinction is between metrics that reflect hashtag effects and metrics that do not.

MetricUseful for hashtag testing?Why
Reach from hashtags (where available)BestDirectly attributes impressions to tag discovery
Total reach / impressionsUsableContaminated by follower reach and recommendation, but the trend is informative
Non-follower reachGoodIsolates discovery, which is what hashtags affect
LikesWeakReflects content quality far more than discovery
Follows from postGood secondaryIndicates you reached the right audience, not just more people
Saves and sharesGood secondaryBest proxies for reaching genuinely interested people

Where the platform exposes a hashtag-specific reach figure in its insights, use it — it is the only metric that isolates the variable directly. Where it does not, non-follower reach is the best available substitute, since hashtag discovery is by definition non-follower discovery.

The minimum viable test

A test that can actually distinguish signal from noise:

  1. Build two tag sets, A and B, that differ in one deliberate way — for example, A weighted broad and B weighted niche. Change one variable, not five.
  2. Post at least five times with each, alternating. Ten posts minimum for the whole test. Fewer than five per variant and post-to-post variance will swamp the effect.
  3. Hold everything else as steady as you reasonably can — similar content type, similar posting times, similar format.
  4. Record non-follower reach and saves for each post in a spreadsheet as you go, not from memory afterwards.
  5. Compare the medians, not the means. One viral post will distort an average completely; the median tells you what a typical post did.
  6. Require a large difference before believing it. If B’s median is 8% higher, that is noise. If it is 40% higher across five posts each, that is worth acting on.

The confounds that ruin most tests

Tests actually worth running

Not all comparisons are equally informative. In rough order of how much you learn:

  1. Broad-weighted versus niche-weighted. The highest-value test, because the answer depends on your account size and nobody else can tell you what it is for you.
  2. Varied per post versus a fixed block. Directly tests the most common piece of advice against your own data.
  3. Eight tags versus three. Tests whether count matters for you at all — frequently the answer is that it barely does, which is useful to know.
  4. Search-shaped tags versus community-label tags. Particularly informative on TikTok and Instagram, where search is a growing share of discovery.
  5. Caption versus first comment. Worth running once so you can stop wondering. Most people find no meaningful difference.

What testing will probably tell you

Two findings recur often enough to be worth predicting. First, the effect size of hashtag changes is usually smaller than expected — typically visible but modest compared with the effect of the content itself. That is genuinely useful information, because it tells you how to allocate your effort.

Second, niche-weighted sets tend to produce fewer impressions but more saves, shares, and follows. If you are measuring reach alone you may conclude niche tags underperform; if you are measuring what compounds, you will often conclude the opposite. Which metric you chose at the start determines the answer, so choose it deliberately.

Keeping records that are worth anything

A spreadsheet with one row per post: date, content type, tag variant, non-follower reach, saves, follows. Six columns, filled in as you post. After thirty posts you have something no generic advice can give you — evidence about your own audience. That is the point at which hashtag decisions stop being guesswork.

Put this into practiceTagly generates ranked, platform-tuned hashtags and keywords for any topic — free, no login, results in seconds.Open the generator →

Frequently asked questions

How many posts do I need to test a hashtag set?

At least five posts per variant, so ten posts minimum for an A/B comparison. Fewer than that and normal post-to-post variance will be larger than any hashtag effect.

Which metric should I use to judge hashtag performance?

Hashtag-specific reach where your platform reports it, since it isolates the variable directly. Otherwise use non-follower reach, which captures discovery. Likes are a poor choice because they reflect content quality more than discovery.

Why should I compare medians instead of averages?

A single viral post distorts an average so badly that it describes none of your posts. The median reflects what typically happens, which is where hashtag effects are visible.

How big a difference is worth acting on?

A large one. An 8% difference in medians is noise. A 40% difference sustained across five posts per variant is a real signal.

What ruins most hashtag tests?

Changing content type at the same time as tags, one viral post skewing averages, follower growth mid-test, testing during an unusual period, and stopping early because early results look favourable.

What do hashtag tests usually reveal?

That hashtag effects are real but smaller than expected relative to content quality, and that niche-weighted sets often produce fewer impressions but more saves, shares, and follows.

Keep reading