Threads or long posts
The post that created instabid.lol was not a thread. It was not a long-form Premium essay either. On 19 August 2026 at 11:08pm, Jonathan Wilke posted a single screenshot plus a link. That post reached 1.6 million views.
The quoted post underneath it was even shorter: “Almost there… Just checking the Polar web hook setup.”
Two 2026 write-ups treat posting format as settled. They do not agree. One says the algorithm used to reward threads for the engagement loop they created, and that in 2026 this flipped: long-form single posts now consistently reach farther than the same material split into a thread. The other cites benchmarks in which threads generate 3x more engagement and 63% more impressions than single tweets.
We have one data point, not a study. This piece lays out the contradiction, says why those benchmarks are usually unfalsifiable, and records what we actually observed.
Two 2026 sources, opposite answers
In May 2026, a Fansgurus growth guide listed long-form posts as a core method. The claim is specific. The algorithm, it says, previously rewarded threads for the engagement loop they created. In 2026 that flipped. Long-form single posts, written up to the extended character limit, “consistently get higher reach than equivalent thread content.”
The same guide gives a mechanic. Threads fragment likes, replies, and dwell time across several posts, each scored on its own. A single long post concentrates those signals. Threads, it adds, still work for some uses. As a default, long-form is “now the dominant choice.”
On 23 May 2026, Farid Sukurov published a Medium write-up on threads, replies, and quote posts. In the thread section he cites “recent benchmarks” claiming threads generate 3x more engagement and 63% more impressions than single tweets. He calls threads “the most reliable format for building authority and driving sustained growth.”
Those two statements cannot both be generally true. If long-form single posts consistently out-reach equivalent threads, the 3x / 63% thread advantage is not a current general fact. If those benchmarks describe the live feed, the 2026 “flip” is not a current general fact.
Neither piece publishes a methods appendix. Fansgurus does not name the accounts, dates, or post pairings behind “consistently.” Sukurov does not name the study behind “recent benchmarks.” The numbers are also not even the same quantity. Reach is not engagement. Impressions are not replies. A thread can win one and lose another.
Later in the same Medium piece, Sukurov writes that article links now get higher reach than equivalent thread content. That does not resolve the earlier benchmark. It adds a third claim in the same article.
Why format benchmarks are usually unfalsifiable
A format claim is easy to print and hard to pin down. The unit is unstable. “Thread” can mean two self-replies or twelve. “Single tweet” can mean 40 characters or 4,000. “Long-form” can mean a Premium expand, an article card, or a 280-character post with a screenshot. If two writers do not lock those definitions, their percentages are not about the same objects.
The comparison is also contaminated by everything that is not format. Account age, follower graph, Premium status, niche, hour of day, whether the first line works as a hook, whether the post contains a link, whether the author replies in the first hour: any of those can dwarf the difference between one post and a chain. A guide that does not hold those constant is describing a mix, not a format effect.
Then the arithmetic hides. Thread impressions can be summed across every post in the chain. If a hook had 10,000 impressions and six follow-ups had 1,000 each, a writer who added them would print 16,000 against one single-post total. Unless the write-up says whether it summed, averaged, or took the hook only, “63% more impressions” is not a number you can audit.
Engagement rates have the same hole. If the denominator is impressions on the hook, a thread looks hotter. If the denominator is impressions on every post, it looks cooler. If replies on post four count as thread engagement, the format is being credited for a conversation that a single post might have produced in one reply well.
Survivorship finishes the job. The threads people remember are the ones that ran. The long posts people remember are the ones that expanded. Failed experiments do not get cited in May roundups. A benchmark built from viral examples will always find that the winning format wins.
None of this proves either source is lying. It means a reader cannot falsify them from the text. You cannot re-run a study that was never specified.
What a longer character limit changed
The structural fact is simpler than the benchmarks. X’s own help page defines a longer post as one that extends past the typical 280-character limit, up to 25,000 characters, as an X Premium feature.
That is not a vibe. It is a change in how many objects a thought has to occupy.
Before longer posts, anything that would not fit in 280 characters had to become a thread or leave the site. A thread is several posts, several timestamps, several early-engagement clocks, several reply wells. The hook is a separate post from the payoff. A reader can like the first and never open the rest. An algorithm can boost the first and starve<<
The quoted post underneath it was even shorter: “Almost there… Just checking the Polar web hook setup.”
Two 2026 write-ups treat posting format as settled. They do not agree. One says the algorithm used to reward threads for the engagement loop they created, and that in 2026 this flipped: long-form single posts now consistently reach farther than the same material split into a thread. The other cites benchmarks in which threads generate 3x more engagement and 63% more impressions than single tweets.
We have one data point, not a study. This piece lays out the contradiction, says why those benchmarks are usually unfalsifiable, and records what we actually observed.
Early scoring can make both stories look true
Buffer’s 2026 account of the X timeline, by Rochi Zalani, does not pick a winner between threads and long posts. It describes a scoring process that can make either format look dominant depending on the first hour.
The post is not shown to everyone at once. Buffer writes that X tests it with a small group, about 5 to 15 percent of followers, and watches the first hour. Replies and saves are treated as strong signals. If they arrive quickly, distribution expands, first to more followers, then to people who do not follow the account. If they do not, distribution slows or stops. The system keeps rechecking. A later repost can start a second life.
Under that rule, a thread can “work” without being intrinsically better. The hook is short. It is easy to reply to. Each follow-up is another chance for a reply-to-reply. Early replies on post one can buy distribution for the chain even if later posts are unread.
A long post can “work” under the same rule for the opposite reason. All the dwell is one object. A save is a save of the whole argument. There is no drop-off between tweets to leak the score. If the first 280 characters earn the expand, the extra time on the post is credited to the only ID that exists.

The only format test that means anything is the one run on your own account.
Buffer also notes that the For You feed and the Following feed are different surfaces. Format talk is almost always about For You. A thread that dominates in a follower timeline can still lose the recommendation contest, and a long post that never leaves the author’s circle can still look fine in native analytics. Two writers sampling different surfaces will publish two climates.
Early-engagement scoring explains how both 2026 claims can be locally true. It does not pick a format. The first hour, the reply well, and the test group can swamp the difference between a chain and a scroll.
One screenshot is not a study
What we observed is narrower than either guide.
The launch post for outbid.lol, which is the post that created instabid.lol as a reaction, went out at 11:08pm on 19 August 2026. It was a screenshot and a link. It reached 1.6 million views. The quoted post under it was a status line about a Polar webhook. There was no numbered chain. There was no 25,000-character body. There was no hook promising five frameworks.
That is one post, by one account, at one hour, with one image and one URL. It is not a controlled comparison of threads against long-form. It does not measure a median. It does not tell you what would have happened if the same screenshot had been tweet one of seven.
It also fails both playbooks in small ways. Fansgurus is arguing for long-form as the default, not for a short caption over a picture. Sukurov is arguing for threads as the growth engine, not for a single live announcement. If you used either article as a checklist, you would not have produced the post we actually watched travel.
We are not going to dress that up as proof that “short wins.” Outliers are how format myths get built. A million-view post is a fact about that post. It is not a rate. It is not a recommendation. The only thing it falsifies is the stronger version of either claim: that you must thread, or that you must write long, or else the post will not move.
We did not A/B the same text as a thread, a long post, and a short image post. We did not log first-hour reply velocity. One data point. Not a study. That sentence is the finding.
How to test format on your own account
If the public numbers cannot be audited, the only comparison that can is yours, and even that is crude.
Write one argument. Keep the first 280 characters identical. Post it three ways on three comparable days: as a thread of the same beats, as a single long post if you have Premium, and as a short post with the rest of the argument in an image or a reply.
Do not change the niche, the hour band, or whether you reply to comments in the first hour. Log impressions, profile visits, replies, bookmarks, and follows at one hour, six hours, and 24 hours. Repeat the set at least five times.
That still will not be a paper. Five triples is not a Fansgurus “consistently” and not a Medium “benchmark.” It is enough to see whether, on your graph, the thread’s summed impressions are an artifact of extra posts, whether the long post’s dwell shows up as bookmarks, and whether the short version dies when there is no expand and no chain.
Watch the denominators. If you add every tweet in the thread, divide by the number of tweets before you call it a win. If a follow happens two hours later, do not assign it to format unless you have no other live posts. If you only reply in the first hour on thread days, you have tested reply behavior, not format.
Do not import someone else’s 3x. Do not import someone else’s flip. Those figures were published without the tables that would let you check them. Your own table will be ugly and small. It will also be about your account, which is the only ranker sample you can see.
If you do not have Premium, you cannot run the long-post cell. Then the test is thread versus short. If you only post from a large account, you will not see what a small-account test group does. The unfalsifiable move is to skip the test and quote a roundup.
The question is still open
We are not going to pick a winner we cannot support.
A closed answer would need a named sample, locked definitions, and a denominator. Fansgurus does not supply that. Sukurov’s cited benchmarks do not supply that. Buffer describes a scoring process, not a format ranking. We have a screenshot, a link, and 1.6 million views.
Until someone publishes the table, “threads or long posts” is not a strategy. It is an argument about a ranker nobody outside the company can inspect, dressed as a percentage. Post the thing you can finish at 11:08pm. Log what it does. Treat May roundups as claims, not as weather. The format question is open.