A small collectible figurine standing before an immense curving wall of glowing cream panels

X decides what you see with Grok now

On 19 August 2026 at 11:08pm, Jonathan Wilke posted a launch for outbid.lol. The post was one screenshot and one link. It was not a thread. It crossed 1.6 million views and more than 800 likes.

instabid.lol exists because that post existed. Within days the mechanic was cloned, including by us. Our board later held 41 rows in total and five visible listings. Reach and outcome were not the same thing.

The system that moved Wilke’s post is not the ranking stack X showed the public three years earlier. In January 2026, according to both Fansgurus and SocialPilot, xAI replaced the legacy recommendation system with a Grok-powered Transformer model that reads every post and makes roughly 5 billion ranking decisions a day. The code was open-sourced on GitHub.

Graham Mann notes that X open-sourced “the algorithm” in March 2023, but that the version running today looks nothing like that release. A 2023 dump is not a map of 2026 distribution.

This is a report on what the replacement changed for a launch post, using Wilke’s post and our own board as the case. It is not a growth guide. It does not claim to have seen the model’s scores.

January 2026 was a replacement, not a ranking tweak

For most of the platform’s life, “the algorithm” was a phrase people used for a bundle of retrieval tricks, engagement graphs, and hand-tuned weights. The March 2023 open-source release let outsiders inspect a version of that bundle. Graham Mann’s point is the one that matters for anyone still quoting that repo: the production ranker is no longer that system.

Fansgurus and SocialPilot both describe the January 2026 change as a swap. Legacy recommendation out. A Grok-powered Transformer in. The new model is described as reading every post, then issuing on the order of 5 billion ranking decisions a day. That is a different object than a feature list with a few weights turned up.

A replacement of that kind changes the unit of judgment. The old public story was that posts were retrieved, filtered, and scored from signals people could name: follow graph, keywords, hashtags, recent likes. The new public story is that a language model looks at the post itself, predicts what a given reader will do, and ranks on that prediction. The GitHub release is the transparency claim attached to that story. It is not, by itself, proof that any particular post was scored a particular way.

For a launch, the practical difference is timing and opacity. You still hit publish once. The system still has to decide, many times per day, whether that post is a candidate for a given timeline. What changed is the thing doing the deciding. If you are posting a product screenshot at 11:08pm, you are no longer trying to match a keyword bucket. You are giving a Transformer a document and asking it to predict engagement.

As of August 2026, SocialPilot’s algorithm recap reported no confirmed further ranking changes. It said the Grok-powered ranking still governed both For You and Following. The standing advice attached to that recap was replies, native video, and consistent posting in one niche. That is the public baseline this article is working from. It is also the ceiling of what outside reporting currently confirms.

What “reads every post” changes versus keyword-era ranking

Keyword-era ranking treated text as an index. Hashtags were a discovery handle. Phrases were features. A post about a leaderboard could be routed to people who had engaged with “leaderboard,” “Instagram,” or adjacent tokens. If you did not use the tokens the retriever expected, you needed the follow graph, a quote, or luck.

A model that SocialPilot says “reads your actual words” is doing something else. It is not waiting for a tag to classify the post. It is building a representation of the post as a whole: what it is about, who it is for, what kind of reply it is likely to attract. SocialPilot also says the model runs sentiment analysis, and that constructive posts get wider distribution while combative bait is throttled even when it generates engagement. That claim belongs to SocialPilot. It is not something we can inspect on a live timeline.

The shift matters for a launch because launches are dense, local documents. Wilke’s post was a picture of a product and a URL. A keyword system would have needed the caption, the link domain, or the author’s existing cluster to know what it was. A reader-model can, in principle, look at the screenshot, the surrounding text, and the likely audience, then place the post next to people already lingering on indie launches, bidding mechanics, or Instagram-growth objects.

“In principle” is doing work there. Fansgurus describes candidate sourcing as a mix of in-network retrieval and out-of-network neural ranking, then a heavy ranker that predicts actions. That is a pipeline story, not a screenshot of Wilke’s scores. What we can say without overreaching is narrower: if the model reads the post, then the post’s meaning is an input, not an afterthought. Hashtag stuffing is no longer the discovery lever it was when the retriever needed handles. SocialPilot states that directly: the system understands the post from the text itself.

The other change is negative space. Keyword ranking failed silently when you missed the index. Semantic ranking can fail in the opposite way. It can decide a post is “about” a cluster you did not intend, or that its tone predicts mute and skip rather than reply. A launch screenshot of a bidding board can be read as a product, a meme, a clone-farm, or spam, depending on what else the model has seen that week. We do not get a label back.

That is the operational fact of “reads every post.” You are no longer stuffing tokens into a hopper. You are submitting a complete object to a ranker that claims to understand it. The object has to be legible in one pass. That will matter again when we get to why a single screenshot traveled and a typical thread did not.

The first hour is a test, not a publication

Buffer describes the distribution mechanic in staged terms. X tests a post with a small group, scores it on early engagement — especially replies and saves — then expands or limits reach accordingly. In Buffer’s account, the first audience is a slice of followers, on the order of 5 to 15 percent, watched in the first hour. Strong early reaction opens a wider follower set, then out-of-network For You. Weak reaction slows the post and can stop it.

That is a test loop, not a broadcast. Hitting publish puts the post into an experiment. The ranker is asking whether this document produces the actions it is trying to predict. If it does, more people see it. If it does not, the experiment ends.

SocialPilot puts the same pressure on the first 30 minutes and calls early engagement velocity the signal that decides whether a post is pushed or left to die. The two write-ups disagree on the exact window. They agree on the shape: early reaction is not a vanity metric. It is the gate.

For a launch at 11:08pm, that gate is awkward. Late posts can still find an awake slice of a global follower graph. They can also test against an empty room and never recover. We do not have Wilke’s first-hour panel. We have the end state: 1.6 million views. Between those two facts sits a sequence of expand-or-kill decisions that are not public.

The test loop also explains why clones of the same mechanic do not automatically inherit the same distribution. Each post is its own experiment. The original is a new object. The fifth screenshot of a bidding board is a different document, scored against a feed that has already seen the first four. Semantic ranking, if it works as described, can tell them apart. It can also collapse them into one cluster and fatigue the reader. From outside, both look like “the same idea.” Inside the ranker, they are separate trials.

Buffer’s account has one more detail that launch posts tend to ignore. Even after a post starts spreading, the system keeps checking performance. A late quote or a second wave of replies can reopen distribution. Silence after a spike can close it. The 1.6 million views are a cumulative result of those checks. They are not a single decision made at publish time.

Replies and saves are the currency

If the first hour is a test, the question is what the test is measuring. Buffer is specific: early engagement, especially replies and saves. Likes are not the headline signal in that description. Conversation and the decision to keep the post are.

SocialPilot is more blunt. It says replies outweigh likes, and that a post with 20 genuine replies beats a post with 100 silent likes. It also says the ranker scores posts on the probability they will spark real conversation. That is the currency claim, attributed to SocialPilot, not to our own logs.

A reply is costly. It requires reading, a decision, and text. A save is costly in a different way: it is a private judgment that the post is worth retrieving. A like is cheap. A ranking system that is trying to predict attention will overweight the costly actions if those actions correlate with more attention later. That is the logic the public write-ups are selling. It is consistent with a model that reads posts rather than counting tokens.

a vinyl figurine diorama of a small robot scoring tweet cards at a desk while two other figurines pass it handwritten reply notes and a bookmark tab

Early replies decide the second wave. The first audience is small on purpose.

For a launch post, the implication is uncomfortable. A screenshot of a live product invites a specific kind of reply: “how does this work,” “I built one too,” “this is a clone,” “here’s the link in a different niche.” Those replies are not community-management extras. Under Buffer’s and SocialPilot’s accounts, they are the fuel the expand step is waiting for.

Author replies sit inside that same loop. Fansgurus describes author-engaged reply chains as the highest-value signal in the system it claims to have read out of the open-source weights. We are not repeating those multipliers here as facts about production. The directional point is enough: if conversation depth is what the ranker wants, leaving a launch post unanswered is a decision to starve the test.

Saves are harder to see from outside. There is no public save count on a post. Buffer still names them as a scoring input. For a product screenshot, a save is a plausible action: someone wants the URL later, wants to copy the layout, wants to show a friend. We cannot know whether Wilke’s post was save-heavy. We can say that a single dense image is the kind of object people save, and a 12-post thread is the kind of object people skim.

Likes remain visible, which is why they still dominate launch recaps. Wilke’s post had 800-plus. That number tells you the post was not ignored. It does not tell you why the ranker kept expanding it. If replies and saves are the real currency, then a likes screenshot is the wrong receipt.

The other currency SocialPilot names in August 2026 is format: native video, and consistent posting in one niche. A one-off launch is, by definition, not a consistency strategy. It can still clear the test if the first panel talks back. It cannot rely on a niche history it does not have.

Why a single screenshot outperformed most threads that week

The week of 19 August was full of threads. Founder threads, teardown threads, “here’s the stack” threads. Wilke posted one image and a link. The image did the explaining. The link was the ask. The post crossed 1.6 million views.

We do not have a controlled comparison against “most threads that week.” We have a visible outlier and a format contrast. The contrast is still worth taking seriously, because it lines up with how the current ranker is described, and because it cuts against a 2023 habit that has not fully died.

A thread asks the model for a sequence of clicks. Each card has to earn the next. Graham Mann writes that short threads with proof still work and that long mega-threads do not. SocialPilot says the same in different words: tight 3–6 post threads, first post as a standalone hook. Wilke skipped the sequence. The screenshot was the hook and the body.

A leaderboard, or a bidding board, is a visual argument. You can see the rows, the names, the mechanic, the joke. A Transformer that reads the post — and, in SocialPilot’s account, reads images as well as text — is being handed a complete object. It does not need seven tweets to know what the product is. Neither does a human scrolling For You at speed.

a collectible figurine holding up one Polaroid of a tiny leaderboard, with a discarded stack of unread thread-card figurines slumped beside it

One screenshot, one link, 1.6 million views. Format was not the variable everyone assumed.

The link is the unresolved part. SocialPilot says posts built around external links get less reach in the main feed, and that native images and video get priority. Wilke’s post had a link and still traveled. Possible explanations, none of which we can confirm: the screenshot carried the post; the link was secondary in the model’s reading; the conversation around the mechanic overwhelmed a link penalty; the 1.6 million views included quote-tweet and reply traffic that the ranker treated as native. Pick one if you want a story. The honest version is that a documented link penalty and a documented 1.6 million-view linked post both exist in the same week.

What the screenshot did that most threads did not is remove delay. There was nothing to unroll. There was something to argue with immediately: the mechanic, the names on the board, the fact that a product this simple was live. Replies could start from the first frame. Under Buffer’s test-and-expand loop, that is the entire game.

The clones tested that theory in public. Within days the mechanic was copied. Each clone was also, in many cases, one screenshot and a link. They did not all become 1.6 million-view posts. Same format, later in the week, into a feed that had already learned the object. If the model reads every post, it can also read repetition. Format was not a cheat code. It was a way to be understood quickly the first time.

Our own board is the drab end of that story. We shipped a clone into a market that had already seen the original. We ended up with 41 rows total and five visible listings. The screenshot format got the idea across. It did not invent demand.

What a launch post is actually asking the model to do

Strip the folklore and a launch post is a ranking request. You are asking the Grok-powered model to do three things in sequence.

First: represent the post. Read the text, the image, the link, the author. Decide what it is. In the keyword era, you helped that step with tags and formula. Now, per SocialPilot, the model claims to get there from the words and the media. A screenshot of a working product is a strong representation. A thread that hides the product until tweet four is a weak one.

Second: test it. Show it to a small group, as Buffer describes, and score the costly actions. Replies. Saves. Then decide whether the predicted engagement is high enough to spend more inventory. A launch that does not invite a reply is asking to fail this step. “We’re live” is not a question. A board people can argue about is.

Third: route it out of network. This is the step founders mean when they say a post “took off.” Fansgurus describes out-of-network ranking as the path by which an account with few followers can still reach strangers. That path is exactly what a 1.6 million-view launch requires. It is also the step you cannot see. You see the view counter. You do not see which SimCluster, which predicted-reply probability, which diversity cap, which follower slice.

The launch-post mistake is to treat step three as a format trick and skip steps one and two. Threads were a 2023 answer to dwell time. In 2026, per the sources above, the ranker already reads the whole post and already overweights conversation. Extra cards are extra chances to drop. Native video is the format SocialPilot still puts first as of August 2026. A still screenshot is not video. It is, however, native, complete, and easy to reply to. That combination is a better match for the described system than a long thread with a link at the end.

There is a second request hiding in a launch post: classify the account. SocialPilot says consistent posting in one niche is still the standing advice, because the ranker routes by recent behavior and topic. A single launch does not build that history. Wilke’s post had to work as a document, not as the tenth post in a niche feed. That is a harder ask. It succeeded once, in public, with a simple image. It is not a template that survives contact with a cloned week.

What nobody can verify from outside

We can count views on a post. We cannot see the ranking decisions.

We do not know how many of Wilke’s 1.6 million views came from For You, from Following, from notifications, from quote tweets, or from people opening the post out of a reply thread. SocialPilot says the Grok-powered ranking governs both For You and Following as of August 2026. That still leaves several surfaces. A view is a view. It is not a map.

We do not have reply counts, save counts, dwell time, or video-equivalent watch data for that post. We have likes, because likes are displayed. If Buffer and SocialPilot are right about replies and saves, then the public receipt is the wrong unit. Anyone recapping the launch from the like count is recapping the metric the ranker is said to value least.

We do not know whether the open-source GitHub code matches production weights on the day Wilke posted. Graham Mann already warns that the 2023 release does not describe the live system. The January 2026 release is a newer claim of transparency. A claim of transparency is not the same as a live score dump for one post at 11:08pm.

We do not know how the model read the screenshot. “Reads every post” is the line Fansgurus and SocialPilot both use. It does not tell us whether the image was treated as a product demo, a meme, or a link wrapper. It does not tell us whether the URL was penalized, ignored, or used as a topic hint.

We do not know why the clones underperformed the original, including ours, beyond the obvious: they were later, they were copies, and the feed had seen the mechanic. That is an outside reading. It is not a ranking log.

We do know what the clone produced on our side. Forty-one rows in total. Five visible listings. That is the part of the story that does not need a Transformer to interpret. Distribution can create a crowd at the post. It cannot fill a board. Treating 1.6 million views as proof of a market was the error available to anyone watching that week. The views were real. The listings were countable. They were not the same measurement.

The last unverifiable is the one people will keep arguing. Did Grok “decide” to boost a simple launch because it was original, native, and easy to talk about? Or did a dense screenshot hit an already-awake builder cluster, pick up replies, and ride Buffer’s expand loop until the view counter looked like destiny? From outside, those two stories are indistinguishable. They use the same public facts. They imply different next posts.

As of August 2026, SocialPilot said the ranking framework had not been replaced again. Replies, native video, one niche. That is the available advice. It is not a reconstruction of 19 August. A launch post still has to survive a test. The model still reads the post. What it thought of that particular screenshot is not a fact we have.