Most polls are a list, and a list is right when the options are ideas. But a good number of the questions a stream actually asks are not ideas at all — they are two things you could show someone. Which thumbnail. Which colourway. Which of these two builds. Which cover art. A list of words describing pictures is a strange way to ask about pictures.
Two, and only two
This-or-that is exactly two options. That is a rule rather than a suggestion, and it is the constraint that makes the format work: two things can share a screen at a size where both are legible, and the result can be one bar rather than a stack of them. A third option turns it back into a list with pictures, which is worse than either.
It also changes what the result means. On a three-option poll a 40% winner is a plurality and the room is split three ways. On a two-option poll, 58% against 42% is a statement — everybody chose, nobody abstained into a middle option, and the gap is the whole story.
The questions it is good at
Anything where the answer is visual and the audience can judge it instantly. Two thumbnails for the video you are about to upload — this is the best use of the format and it is genuinely useful research, because the people voting are exactly the people who would or would not click. Two character builds. Two loadouts. Two possible overlays. Two cover images for the same track.
It is also good at questions that are not visual at all but are binary and a bit silly, in which case the pictures do a different job: they set the tone. "Ranked, no talking" against "chaos lobby" reads differently with a photo attached to each.
The images are optional, and sometimes better left out
You can run a this-or-that with no images at all, and it renders as a two-option poll with big tiles. That is a perfectly good thing to run, and it is worth knowing because the alternative — feeling obliged to find a picture before you can ask a two-option question — is exactly the friction that stops a poll happening at all.
When you do use images, use ones that are legible small. The tile on stream is not full-screen, and a busy screenshot with detail in the corners reads as a coloured rectangle. A face, a silhouette, a strong shape or a single object survives the shrink; a wide landscape does not.
Say what the difference is
The commonest mistake is showing two images that differ in a way nobody can see at that size and asking which is better. If the two thumbnails differ only in the font weight of the title, the poll is measuring noise, and worse, it will produce a confident-looking result you might act on.
Label them. The words under each tile are not decoration — they tell the room what they are being asked to compare. "Face, big text" against "gameplay, no text" is a question. "A" against "B" is a coin flip you have dressed up.
Use the result honestly
A poll of your own viewers is a sample of your own viewers, which is the right sample for "what should this stream do next" and the wrong one for "what will strangers click on". Your regulars know your face and your in-jokes; the people your thumbnail has to reach do not. A 60/40 result from your chat is a useful nudge and not a verdict.
And keep it short, the way any poll should be. This-or-that in particular resolves fast — people do not deliberate over two pictures — so ninety seconds is generous and thirty is often enough.
Two rounds beat one long list
The format also gives you a way to handle more than two candidates without a list: run it as a bracket. Four thumbnails become two head-to-heads and then a final, which takes about four minutes and is considerably more entertaining than a single poll with four options, because each round has a loser and people have opinions about losers.
It produces a better answer too. In a four-option list the options at the bottom lose votes purely for being at the bottom. In a bracket every pairing is judged on its own, and the thing that wins has actually beaten something rather than merely out-placed it.
It is a segment, not just a widget
The reason to bother with pictures at all is that a visual choice is watchable. A list of words on screen is a mechanism; two images with a bar sliding between them is a small piece of television, and it holds attention for the thirty seconds it needs to. That is why it works as a between-games beat rather than only as a decision tool.
It also survives being clipped, which the platform-native polls do not — they render in each viewer’s own interface and vanish from the recording. An overlay this-or-that is part of the video, so the moment the room chose the ridiculous option is still there in the VOD, with the bar and the numbers.
Where it sits among the rest
It is one of five poll types, and the one to reach for whenever the answer is a thing rather than an idea. A list handles the ideas, a multi-select handles a shortlist, a quiz asks a question with a right answer, and a word cloud asks the room for a word instead of choosing from yours. This-or-that is the one that shows rather than tells, which is a strange thing for a stream not to do.