RC RANDOM CHAOS

Perplexity cites 215,128 template pages as evidence

Three sites generated 215,128 scripted "best software" pages that AI engines like Perplexity cite as sources - fueling misinformation, SEO abuse, and malware.

· 7 min read
Perplexity cites 215,128 template pages as evidence

Three websites published 215,128 pages ranking the “best software” for tasks almost nobody ever typed into a search bar by hand. The pages weren’t written for readers. They were assembled for retrieval - for the exact moment an AI assistant needs a source to hang an answer on. Perplexity, the AI search engine that sells itself on showing its citations, pulls from these pages and lists them as evidence. The citation is the product. The content behind it ranges from arbitrary to actively dangerous.

This isn’t a story about low-quality blog posts. It’s a story about what happens when the machine that answers your questions treats “this URL exists and matches the query” as a stand-in for “this is true.”

How three sites become 215,128 pages

Nobody writes 215,128 articles. You write one template and pour data into it.

The technique is called programmatic SEO, and it’s decades old. You take a page skeleton - “The 7 Best [TOOL TYPE] for [USE CASE] in 2026” - and cross-multiply two lists. Fifty software categories times a few hundred use cases times a handful of audience modifiers, and you’ve generated a six-figure page count from an afternoon of setup. Each page gets a unique URL, a unique title, a filled-in intro, a ranked list, and a comparison table. Then you drop every URL into an XML sitemap and hand it to the search crawlers on a plate.

The economics only work at this scale because the marginal cost of page number 200,000 is close to zero. No writer touched it. No editor reviewed it. The “rankings” inside were not produced by installing the software and testing it. They were produced by a script, and the order is frequently set by which vendor pays the largest affiliate commission, not which tool works.

That’s the thing to hold onto: these aren’t opinions that happen to be wrong. They’re outputs of a spreadsheet dressed up as editorial judgment.

Why an AI search engine reaches for them anyway

Perplexity and tools like it don’t “know” answers. They retrieve them. Ask for the best free PDF editor and the system runs a search, grabs a handful of top-matching pages, feeds that text into a language model, and asks the model to summarize with citations. This is retrieval-augmented generation, and its blind spot is structural: the retrieval step optimizes for relevance and freshness, not for whether the source did any real work.

A programmatically generated page is engineered to win that retrieval step. It matches the query almost word for word, because the query was reverse-engineered into the page title. It’s recently published, because the whole farm can be regenerated on a schedule. It has clean headings, a comparison table, and structured markup - exactly the shape a machine parses easily. To a retrieval system, a page built by a script to look authoritative and a page built by a practitioner who tested twelve tools are nearly indistinguishable. Both are just text with a matching URL.

So the AI cites the farm. And here’s the compounding problem: once the model repeats the farm’s ranking, that ranking gains a new kind of credibility. A user sees it delivered in calm, sourced prose. The affiliate spreadsheet has been laundered into an answer.

The misinformation layer is the boring part

Start with the least alarming failure, because it’s the most common. The lists are simply not grounded in anything.

When a page claims a tool is the “best for small teams,” ask what test produced that claim. On these pages, the answer is usually none. The ordering reflects commission rates, existing brand deals, or whatever the generation script defaulted to. A genuinely good, free, open-source tool with no affiliate program is invisible to this system by design - there’s no payout to rank it, so it doesn’t appear. The reader gets a filtered view of the market that has been quietly shaped by who pays, and the AI passes that filter along without flagging it.

Now add the feedback loop. When enough farms cite each other, and AI answers cite the farms, and new content is written using those AI answers, the same unverified ranking circulates until it looks like consensus. Nobody checked it at any step. It became “true” through repetition. This is how a factual error or a paid placement hardens into something people quote as established, and it’s a failure mode that gets worse as more of the web’s writing is itself machine-generated from prior machine output.

Wrong recommendations waste time and money. Annoying, not catastrophic. The catastrophic version is next.

The part that actually gets you owned

“Best free software” searches have been a top malware delivery channel for years, long before AI search existed. The pattern has a name in the security world: SEO poisoning. Attackers build pages that rank for high-intent download queries - “free video converter,” “best PDF tool,” “download [popular app] free” - and the download button hands you a trojan, an info-stealer, or a fake installer. Campaigns like Gootloader and SolarMarker ran this playbook for years, buying or gaming the exact same “best software” real estate these farms now occupy at industrial scale.

A mass-produced page factory is a gift to that kind of attacker for two reasons.

First, volume and automation are the whole point of both operations. A system that can generate 215,128 pages can generate 215,128 pages with poisoned outbound links, or be compromised and repurposed. The affiliate links, the “official download” buttons, the redirect chains - every one is a place to swap a legitimate destination for a malicious one, and at this scale no human is watching any single page. Typosquatted download domains and lookalike installers fit into this structure without anyone noticing.

Second, and this is the specific new damage, the AI strips out the warning signs a careful human uses. When you land on a sketchy “top 10 downloads” site yourself, you get signals: an ugly layout, aggressive ads, a URL you don’t recognize, that instinct that says close this tab. When Perplexity summarizes the same page, those signals are gone. You get a clean sentence - “a popular option is X, available here” - with a tidy citation. The interface that was supposed to make search safer has removed the friction that kept people cautious. A phishing page that a human might have distrusted on sight arrives pre-endorsed by a tool the user trusts.

That’s the escalation. Old attack, new distribution channel, and the new channel launders the attacker’s credibility for them.

How to read a machine that cites its sources

Citations are not verification. This is the single most useful thing to internalize. A citation tells you the model found a page that contained the claim. It tells you nothing about whether that page tested anything, whether it’s paid placement, or whether it’s hostile. “Sourced” and “checked” are different words for a reason.

A few concrete habits that cost almost nothing:

  • When an AI answer recommends software, click through to the citation before you trust the ranking. Look at who runs the site. A page with no named author, no methodology, and a comparison table stuffed with affiliate links is a farm, not a review. Treat its ranking as an advertisement, because that’s what it is.

  • Never download software from a link inside an AI answer or a “best of” list. Go to the vendor’s official site directly by typing the name into a fresh search and checking the domain, or use your operating system’s trusted package manager or app store. The download destination is exactly where the swap happens, so it’s the one link you should never take on faith.

  • Cross-check the recommendation against a source that has a reputation to lose - a named journalist, an established publication, a maintainer’s own documentation, a community forum where real users argue. If the only places backing a claim are interchangeable “top 10” pages, you don’t have three sources. You have one pattern, printed 215,128 times.

  • For anything that touches money, credentials, or a company machine, assume the recommendation is untested until you’ve confirmed the tool independently. The convenience of an instant answer is not worth handing your endpoint to whoever gamed the retrieval index that week.

The deeper issue sits with the AI companies, and it won’t be fixed by users clicking carefully. A search product that ranks a scripted page farm alongside genuine testing has a sourcing problem it has chosen not to solve, because solving it means spending money to evaluate source quality instead of just source relevance. Until that changes, the citation under an AI answer is a starting point for your own checking, not the end of it. The machine found a page. Whether that page is telling you the truth, selling you a placement, or handing you malware is still your job to determine - and right now, it’s a job the tool is quietly handing back to you while implying it already did the work.

Share

Keep Reading

Stay in the loop

New writing delivered when it's ready. No schedule, no spam.