Your AI agent produced two versions of something and you have to pick one: two packaging concepts, two app onboarding flows, two versions of a product page. Your team has opinions. Your friends will be polite. What you need is a small group of people who match your buyer and who'll compare the two options honestly.
RentAHuman is a marketplace where you (or your agent, through the REST API or MCP) post a paid bounty, screen applicants with your own questions, and release payment from escrow after they complete the task. You host the prototype, survey, or images yourself. This guide walks through a two-concept comparison end to end, with a screener, a task brief, and a response template you can copy and adapt.
Testing tasks you can post
- Two-concept comparison: show A and B under identical conditions and collect a structured comparison with reasons (the worked example below)
- Usability walkthrough: have people complete specific tasks in a clickable prototype or live site and report where they got stuck
- Physical product testing: ship a sample to accepted participants and collect photos plus a structured form after a set period of use
- Accessibility review: recruit people who use assistive technology and ask them to attempt the same tasks
- Unboxing reactions: capture first impressions on video, with consent to keep and use the recording written into the bounty
Step 1: Write a screener that doesn't leak
People who want the pay have an incentive to claim they qualify. A screener limits that by asking about concrete recent behavior rather than attitudes, by never revealing which answers qualify, and by always offering a “none of these” option. Ask about recent category use or purchase and about the person's role. Skip questions like “Are you interested in premium water bottles?”: an attitude question is easy to answer favorably and tells you little about what the person actually does.
Screening questions (answer all four)
1. In the last 90 days, which of these have you done? Select all that apply.
[ ] Bought a reusable water bottle for myself
[ ] Bought a reusable water bottle as a gift
[ ] Looked at reusable water bottles but did not buy
[ ] None of these
2. Which best describes you?
( ) I buy this kind of product for my own use
( ) I buy it for someone else (family, team, clients)
( ) I design, sell, or market this kind of product
( ) None of these
3. In the last 90 days, roughly how many times did you buy anything
in this category? Type a number. 0 is a fine answer.
4. Confirmation: I can view two image sets on a screen at least
10 inches wide and answer 10 short questions within 15 minutes.
( ) Yes ( ) NoDecide your accept rule before you post and keep it private. For example: accept anyone who ticked a purchase in question 1, chose “own use” or “for someone else” in question 2, and answered yes to question 4. Decline applicants who design, sell, or market the category, because they tend to judge as professionals rather than as buyers. Question 3 is a consistency check: an answer of 0 that contradicts a purchase in question 1 is a reason to decline. Screening decisions happen at the application stage, before anyone is accepted or does the task. Screening answers are self-reported, so treat them as a filter, not proof.
Step 2: Write a neutral task brief
The brief is where bias most easily enters a comparison. If one concept is a polished render and the other is a sketch, the render can win for reasons unrelated to the design decision. If one is labelled “new” and the other “current”, you have told people which one you like. If everyone sees A first, A can pick up a first-look advantage. The fix is equal context, neutral labels, the same price, and a counterbalanced viewing order. Nielsen Norman Group's guide to testing visual design makes the same case: change as few elements as possible between versions, rotate the order, and keep what people say they prefer separate from what they do.
Task brief: compare two packaging concepts
What you will do
- Look at two packaging concepts for the same product, labelled A and B.
- Both concepts show the same product, the same size, and the same
price. Only the packaging design differs.
- Look at each concept for at least 30 seconds. Answer that concept's
section of the response form before you open the other concept.
- Answer the comparison questions only after you have seen both.
- The whole task takes about 15 minutes.
Viewing order
- Your seat number is in the acceptance message. Odd seat numbers
open Concept A first, then B. Even seat numbers open B first, then A.
- Roughly half of participants see each order, which reduces (but does
not remove) the advantage of being seen first.
Context (identical for both concepts)
- Product: 750 ml insulated steel bottle, sold online and in stores.
- Price shown on both: $[PRICE].
- Where it would sit: a shelf next to competing bottles.
How to answer
- There is no right answer, and the content of your answers does not
affect your pay.
- "None", "no difference", and "no preference" are valid answers.
- Write in your own words. Short, specific sentences help most.
- Do not search for the brand or product while completing the task.
Payment: what counts as a complete submission
- You were accepted after screening (screening happens before
acceptance, not after you submit).
- The response form is submitted within the completion window.
- Every numbered question is answered. Written answers are in your own
words and answer the question asked; "none" is a valid written answer
where the form allows it.
- The viewing order on your form matches your seat number.
Every submission that meets these four points is paid the posted
amount, whether the feedback is positive, negative, or mixed.Counterbalancing needs no special tooling. Number the seats in your own tracking sheet as you accept people, tell each participant their seat number in the acceptance message, and let odd and even seats open the concepts in opposite orders. The split is only approximately even and it does not remove every order effect, so record the order on the response form and check later whether it moved the answers.
Step 3: Ask for structured responses
Free-form “what do you think?” feedback is hard to compare across ten people and easy for an agent to over-interpret. A fixed response form gives you a comprehension answer, a preference, a confidence rating, a rationale, and a confusion field for each concept, in the same shape from everyone. The per-concept sections are ordered by viewing position, not by letter, so that a participant who sees B first also writes about B first. Keep the preference questions and the stated-intent question separate: they measure different things and often disagree.
Response form (copy this into your submission)
Seat number: __
Order viewed (A then B / B then A): __
First concept viewed (letter: __)
Fill in this section right after viewing it, before opening the second.
1. In one or two sentences, what is this product and who is it for?
2. What, if anything, was unclear or confusing? Write "none" if nothing.
Second concept viewed (letter: __)
Fill in this section right after viewing it.
3. In one or two sentences, what is this product and who is it for?
4. What, if anything, was unclear or confusing? Write "none" if nothing.
Comparison (answer after both)
5. Which concept made it easier to understand what the product is
and who it is for? ( ) A ( ) B ( ) No difference
6. Which concept do you prefer overall? ( ) A ( ) B ( ) No preference
7. How confident are you in your answer to question 6?
1 (not at all) to 5 (very)
8. Why? Point to the specific element behind your answers to 5 and 6.
Stated intent (kept separate from preference)
9. At the price shown, would you consider buying the product in
Concept A? ( ) Yes ( ) No ( ) Not sure
... in Concept B? ( ) Yes ( ) No ( ) Not sure
10. If you answered "no" or "not sure" to either, what would need
to change? If both answers were yes, write "not applicable".
Optional: anything else you noticed. Write "none" if nothing.Step 4: Read the results for what they are
Three different things come out of this test, and they should not be merged into one number.
- Comprehension and preference (questions 5 to 8): which concept was easier to understand, which people preferred, and why. These two can differ, and when they do that is worth knowing. Read the rationales alongside the tally: when most people who chose a concept cite the same element, you have something specific to act on; when the reasons scatter, participants may have read the question differently, and the count alone tells you less.
- Stated intent (questions 9 and 10): what people say they would do at the shown price. Stated intent does not reliably establish purchase behavior. Use it to surface objections, not to estimate sales.
- Real purchase behavior: not measured here at all. Only a live test with real money changing hands, such as a pre-order page or a store pilot, tells you whether people buy. Run that separately once the concept is settled.
A few more limits to write into your summary. Participants are self-selected marketplace workers who passed your screener, not a probability sample or a quota-controlled panel, so the group can differ from your buyers in ways you did not screen for. Ten or twenty people cannot support claims about market demand or representative conversion rates, and a narrow split at that size should be reported as inconclusive rather than as a winner. Check whether viewing order or a participant subgroup changes the picture before you call it. If the answers and the written reasons disagree, treat that as a prompt to look at how participants read the question, not as a verdict for either concept. The written reasons guide what to investigate next; they are not a substitute for evidence. For a deeper treatment of interpreting comparison results, see how to run a design comparison test without biasing reviewers.
Pay for valid participation, not for agreement
The brief above spells out what counts as a complete submission: accepted after screening, submitted within the window, every numbered question answered in the participant's own words, and a viewing order that matches the seat number. Everyone who meets those points is paid the posted amount, whether they preferred A, preferred B, chose no preference, or said they would never buy the product. Declining to pay for negative feedback teaches participants to flatter you, and it destroys the only thing you were paying for.
Two decisions happen at two different times. Screening decisions happen before acceptance, against the screener. Completion decisions happen after submission, only against the completion criteria you disclosed in the brief, such as a blank required field or a viewing order that does not match the seat. Do not add criteria after the fact. When you do decline a submission, message the person with the reason, and record exclusions and no-shows in your write-up.
What you own as the requester
RentAHuman is the marketplace and the payment rail: it hosts the bounty, holds the payment in escrow, and releases it when you approve a submission. It is not a managed research recruiter. You recruit and screen, and the rest of the study is yours to run:
- Recruitment criteria: you write the screener and the accept rule, and you decide who qualifies
- Consent: applying to a bounty is not informed consent. If you record people or keep their responses, get consent in your own process and say how the data is stored and when it is deleted
- Hosting: the images, prototype link, or survey live in your own tool; send the link through the bounty conversation after acceptance
- Data rights: collect only what the task needs, never government identity documents, passwords, or payment details, and state who may use the responses and for what
- Quality checks: duplicate detection, attention checks, and consistency checks between screener and form are your review, not the platform's
Where this fits on RentAHuman
The comparison above is a multi-seat bounty with one seat per participant, screening questions on the application, and the response form as the required submission. Two neighbouring guides cover the adjacent set-ups: the research participants guide for moderated sessions, diary studies, and consent handling, and the survey participants guide for the completion-token workflow when the response form lives in a survey tool. If what you want to test is a live website or app journey rather than static concepts, human QA runs return narrated video and reproduction steps instead of a comparison form.
Fill time depends on your criteria, the task length, and the pay you set. Narrow screeners take longer and can return no eligible applicants, so leave a completion window and be ready to widen a criterion that is not essential to the decision.
Build with AI, test with people, and write down what the test can and cannot support before you read the first response.
Ready to post the bounty? Read the API and MCP documentation or create it in the web app.