Methodology

How we test

How PromptPlay ranks AI story, chat and companion apps: receipts for every claim, a 90-day re-review rule, the Slop Report rubric, and how our own copy is checked.

Rankings

How rankings work

Each ranked list answers one search a reader actually makes, like "Character.AI alternatives". Apps are ordered by how well they fix the problem that sent the reader looking.

When another app wins a category, the list names it in an award. Every app gets watch-outs, including the one ranked first. Prices are written as true "at this review", and the review date is on the page.

Receipts

Receipts and freshness

Every claim about an app comes from a receipt: the source text copied word for word, the link, the kind of source and the day we read it. Receipts are printed at the bottom of each review.

Official sites and app store listings back prices, limits, age ratings and content policies. Reviewer and forum sources only ever appear as what reviewers or users report.

Every review and ranking is checked against the live product at least every 90 days. The page shows when it was last updated, and our content check flags any page that goes past 90 days until someone re-reviews it.

Rubric

The Slop Report rubric

A Slop Report scores up to 6 axes from 0 to 100, lower is better. Each axis has five bands that line up with the Slop Meter tiers, and a score has to sit in the band its receipts support. Axes that do not apply are left out, so an app with no art gets no art score. The overall score is the rounded average of the axes scored.

What each Slop Meter tier means on each axis
TierProseDoes it write like a person?MemoryDoes it remember you?CharactersDoes everyone sound the same?StoryDoes anything happen?PricingDoes it charge you for caring?ArtDoes the art look like a person made it?
Artisanal0-15Specific, varied writing. Stock phrases are rare, and people write or edit what you read.Remembers specific events across sessions, by design.Distinct, consistent voices written or edited by people.Authored plots with stakes, branches and endings.Free or one flat price, with no limit that interrupts a scene.Hand-drawn or finished art, consistent across every image. No generator tells.
Home-cooked16-35Mostly specific, with a few repeated habits.Long-term recall holds for the details that matter, with occasional misses.Consistent characters, with some sameness between them.Structured scenarios or quests that move forward.Clear prices shown before you pay, and limits you can predict.Generated or assisted art that someone fixed and finished: consistent faces, clean hands, deliberate style.
Microwaved36-60Stock phrasing in most replies. Readable, but you have seen it before.Remembers within a session; long-term recall is patchy or up to you.Voices blur together, and personality depends on the character card.A plot exists if you drive it; the app follows more than it leads.Features you will want sit behind a subscription or a currency, with caps you will notice.Clearly generated, with the default gloss and occasional errors, but characters stay recognizable.
Cafeteria61-80A cliche in nearly every reply: whispers, smirks, racing hearts.Forgets between sessions unless you maintain notes yourself.Most characters share one default voice.Scenes loop and stall without heavy steering.Caps or paywalls that land mid-conversation, or a currency that hides the real price.Frequent tells: broken hands or anatomy, garbled text, the same waxy face, or a character who looks different in every picture.
Pure Slop81-100Wall-to-wall stock phrases and stock names. Any two replies could swap places.No memory past the current chat, or it confidently remembers things that never happened.Every character is the same assistant in a different costume.Nothing happens unless you write it.Costs climb as you get attached, or cancelling is a fight.Raw generator output shipped as final: melted features, gibberish text, and no two images of a character match.

What the art axis looks for

We check each image for these tells, in this order.

  1. 1. Anatomy and structure errors

    Hands with extra or fused fingers, melted features, limbs that merge into furniture, backgrounds that stop making sense.

  2. 2. Gibberish text

    Signs, labels, books or UI inside an image with letters that spell nothing.

  3. 3. Generic gloss

    Waxy or airbrushed skin, a yellow or orange-teal cast, flat uniform lighting, the default AI anime look.

  4. 4. Same face and drift

    Every character shares one face, or one character's face changes between images, or the avatar does not match the selfies.

  5. 5. Zero iteration

    Raw generator output shipped as final, with no sign anyone fixed or finished it.

The Fun Meter

Fun runs the other way, 0 to 10, higher is better. It scores whether you would keep playing once the novelty wears off.

Fun Meter bands, 0 to 10, higher is better
ScoreBand
9-10Can't put it down
7-8A good time
5-6Decent
3-4Meh
0-2A chore
Scores

How scores are made

Every score starts with the standard test. People run the same set of test chats with each app and keep the transcripts.

Our slop counter scores the transcripts: stock phrases, repeated lines, forgotten details, scenes that loop and characters who share one voice.

An editor then applies the rubric to what a transcript cannot show, like pricing, and writes one line on each score saying why it sits in its band.

Every score is checked again at each review, at least every 90 days.

Our copy

How we check our own copy

Every page on this site goes through the same slop counter before it ships. The build renders every page, runs the counter over the HTML, and fails if it finds stock phrases, em dashes, curly quotes or anything else on the list. A page you can read here passed.

The example replies we wrote to show what slop looks like are labeled as ours, and the counter skips them.

Slop check: passedThis page went through our slop counter before publishing. How we check our own copy