Researchers Endeavoring To Conduct An Online Study Should Consider

10 min read

Researchers endeavoring to conduct an online study should consider one thing before they even open a browser tab: the internet is not a lab.

That sounds obvious. But most of us forget it the moment we start designing. Because of that, we treat Qualtrics or Prolific or MTurk like a controlled environment. It isn't. Your participants are on couches, in coffee shops, on buses, sharing a laptop with a roommate, answering on a cracked phone screen while a toddler climbs their leg. You don't control the lighting. You don't control the noise. You don't even control whether they're actually reading the questions.

And yet — online studies are where the field is moving. Fast. So if you're going to do this, you need to do it with your eyes open.

What Is an Online Study, Really?

An online study is any research where data collection happens remotely via the internet. This leads to longitudinal diary studies. The method varies. Unmoderated usability tests. Surveys. Card sorts. Also, eye tracking via webcam. Practically speaking, experiments. The constant is distance.

It's Not Just "Surveys on the Internet"

People hear "online study" and picture a Google Form. That's the shallow end. Real online research now includes:

  • Behavioral experiments with reaction-time measures (yes, they work — with caveats)
  • Experience sampling where participants get pinged five times a day for two weeks
  • Remote physiological data via wearable integration
  • Video-based interviews with automated transcription and coding
  • Massive multiplayer paradigms where hundreds of participants interact in real time

The tooling has caught up to the ambition. But the methodology? Still catching up.

Why This Matters More Than You Think

Ten years ago, online data was "convenience sampling" — a necessary evil for underfunded grad students. Now? Top journals publish online experiments regularly. Grant panels expect digital recruitment plans. The pandemic didn't just accelerate this; it legitimized it.

But legitimacy brings scrutiny.

Reviewers know the failure modes. They'll ask about bot detection. Because of that, they'll ask how you verified attention. They'll ask whether your sample represents anything beyond "people who sign up for survey panels." If you don't have answers, your study doesn't get published — or worse, gets published and then torn apart in a replication attempt.

And there's the ethical dimension. In practice, online participants are more vulnerable in some ways, not less. They're isolated. Which means they can't ask a researcher "wait, what does this mean? " in real time. They're often doing this for pocket money, not scientific altruism. That changes the power dynamic.

Easier said than done, but still worth knowing That's the part that actually makes a difference..

How to Design an Online Study That Actually Works

Platform Choice Is a Research Decision, Not a Logistics Decision

Don't pick a platform because it's cheap. Pick it because it fits your design That's the part that actually makes a difference..

Prolific — best for behavioral experiments, high-quality attention, fair pay culture. But: smaller pool, UK/US skew, no built-in experimental engine Simple, but easy to overlook..

MTurk — massive, fast, cheap. But: bot farms, "professional" workers who game attention checks, declining data quality since 2018. Still usable if you layer aggressive quality controls.

CloudResearch / TurkPrime — adds verification layers on top of MTurk. Worth the markup for most studies Not complicated — just consistent..

Qualtrics Panels / Dynata / Lucid — representative sampling by quota. Expensive. Good for population-level claims. Bad for experimental control.

Gorilla / Lab.js / jsPsych / PsychoJS — these are engines, not panels. You bring your own participants. Maximum experimental control. Steep learning curve Small thing, real impact. Practical, not theoretical..

Your own recruitment — email lists, social media, community partnerships. Total control over who sees the link. Total responsibility for ethics, bias, and follow-up Most people skip this — try not to..

The platform shapes your sample. Your sample shapes your conclusions. Choose deliberately.

Pilot Like Your Career Depends On It (It Does)

Run a pilot. " Run a real pilot on your actual platform with your actual target sample. Think about it: not "send it to three lab mates. N = 30–50 minimum Not complicated — just consistent. Took long enough..

What you're looking for:

  • Completion time — if it's 22 minutes but you told them 15, you've broken trust
  • Drop-off points — where do people quit? That's your design flaw, not their fault
  • Attention check failure rates — if 40% fail, your checks are broken or your participants are bots
  • Variance — is everyone scoring 95% on your manipulation check? Your manipulation failed
  • Technical bugs — mobile layout broken? Video won't load on Safari? Timer drift on Firefox?

Fix it. Then pilot again It's one of those things that adds up..

I've seen studies with 800 participants where the randomization logic was inverted for half the sample. Worth adding: eight hundred people. Wasted. Because nobody clicked "preview" on mobile.

Attention Checks: The Art of Not Being Obvious

Standard attention checks ("Select 'Strongly Agree' for this item") catch only the laziest bots and speeders. Sophisticated participants — and sophisticated bots — pass them instantly Small thing, real impact..

Better approaches:

  • Instructional manipulation checks (IMCs) — embed a real instruction inside a block of text: "To show you're reading, please select 'Neither agree nor disagree' for this item only." Looks like a normal item. Isn't.
  • Consistency checks — ask the same question twice, separated by 20 items. Reverse-code one. Flag discrepancies.
  • Free-text traps — "In one sentence, describe what you had for breakfast." Bots hallucinate. Humans write "toast" or "nothing, I'm fasting."
  • Timing traps — hidden timestamp on page load vs. submit. If they "read" a 500-word vignette in 3 seconds, they didn't.
  • Honeypot fields — invisible form fields (CSS display: none). Humans don't see them. Bots fill them.

Layer these. Still, don't rely on one. And pre-register your exclusion criteria. No post-hoc "I'll drop the bottom 15% on attention.

Sampling Bias Is Not a Footnote

Your sample is not "the internet." It's "people on Platform X who saw your study, met your filters, consented, started, and finished."

Every step filters.

  • Platform filter — Prolific users are younger, more educated, more liberal than MTurk. Both are WEIRD (Western, Educated, Industrialized, Rich, Democratic).
  • Filter filter — you screen for "fluent English speakers aged 18–35." You just excluded immigrants, older adults, non-native speakers.
  • Self-selection filter — people who choose your study differ on motivation, curiosity, need for cognition.
  • Completion filter — the 23% who drop out? They're systematically different. More distracted? Less interested? Lower working memory?

You cannot eliminate this. Report your funnel: invited → consented → started → passed attention → completed. Plus, you must document it. Compare demographics at each stage if you can.

And please — stop claiming "representative sample" because you hit census quotas on age and gender. Worth adding: that's not representation. That's quota matching on two variables That's the whole idea..

Mobile Is Not an Afterthought

Over 60% of Prolific sessions are mobile. On the flip side, on some panels, it's 80%. If your study breaks on a phone, you've excluded the majority The details matter here..

Test on:

  • iPhone Safari

  • Android Chrome (multiple versions — Samsung, Pixel, budget devices)

  • iPad Safari (split-screen, keyboard attached, keyboard detached)

  • Mobile Firefox, Edge, Brave — they render differently

Check:

  • Touch targets — 44×44px minimum. Test with predictive text on and off.
  • Auto-advance — don't. Modals trap focus.
  • Viewport meta tag<meta name="viewport" content="width=device-width, initial-scale=1"> or your desktop layout shrinks to unreadable. And the text input? Worth adding: fixed headers jump. Mobile users scroll past the submit button. - Keyboard occlusion — does the virtual keyboard cover the "Next" button? Your 20px radio buttons are unusable. On top of that, - Scroll behavior — iOS Safari bounces. Android doesn't. Put it at the bottom and sticky.

If you're using Qualtrics, enable "Mobile Friendly" and still test manually. Their "mobile preview" lies No workaround needed..

Compensation Ethics: Time Is Not Money

"$10/hour" means nothing if your study takes 45 minutes on mobile but 12 on desktop.

  • Pilot on each device class. Pay pilot participants full rate. They're doing QA work.
  • Set a ceiling. "Maximum 20 minutes" with a soft timeout at 25. Auto-submit with partial data saved. Pay prorated.
  • Reject the "bonus for speed" model. It incentivizes satisficing. Pay fairly for attention, not haste.
  • Transparent timing. Show a progress bar with time elapsed, not just %. Participants pace better.

And if your IRB says "we don't regulate pay rates" — that's a failure of ethics review, not a green light.

Data Quality Theater vs. Actual Quality

Everyone talks about data quality. Few measure it The details matter here..

Theater:

  • "We used three attention checks." (All on page 2. Passed by bots.)
  • "We excluded straight-liners." (On a 7-point scale with 30 items? That's fatigue, not fraud.)
  • "We required 90% approval rate." (On Prolific, that's the default. Meaningless.)

Actual quality signals:

  • Response time distributions — bimodal = two populations (readers vs. clickers). Plot it.
  • Entropy scores — low entropy on Likert scales = straight-lining or patterning. High entropy on free text = engagement.
  • Construct-relevant variance — do your manipulation checks actually move the DV? If not, your manipulation failed or your sample didn't engage.
  • Test-retest reliability — embed 5 repeated items at the end. Correlate. r < .70? Flag the participant.

Publish your quality metrics. Make them a norm.

The Replication Crisis Starts in Your Survey Flow

You didn't randomize block order. You didn't counterbalance scale direction. Which means your manipulation check was the dependent variable. Your "exploratory" analyses are the ones you report And that's really what it comes down to..

Fix it:

  • Pre-register. Not "I have a plan." A timestamped, public, falsifiable plan. Also, asPredicted takes 10 minutes. Plus, - **Randomize everything. Consider this: ** Question order, block order, scale orientation, stimulus assignment. But use your platform's API or Qualtrics' randomizer. Here's the thing — - **Separate manipulation checks from DVs. ** Different pages. Different scale formats. Here's the thing — no carryover. - Label exploratory analyses. "Exploratory" ≠ "unplanned but pretty.In practice, " It means no inference. Because of that, report effect sizes with CIs. Even so, no p-values. - **Share materials.Which means ** Full survey flow. Stimuli. Analysis code. Now, raw data (de-identified). If you can't share, you can't replicate.

A Checklist Before You Launch

Item Done?
Tested on 3+ mobile devices/browsers
Pilot N ≥ 20, paid full rate, feedback collected
Attention checks layered (IMC + consistency + timing + honeypot)
Exclusion criteria pre-registered
Funnel tracking: invited → consented → started → passed → completed
Compensation prorated, ceiling set, no speed bonuses
Randomization implemented (blocks, scales, stimuli)
Manipulation checks separated from DVs
Pre-registration public and timestamped
Materials, code, data sharing plan ready

If any box is unchecked, you're not ready. The 45 minutes to fix it now saves 450 participants' wasted time later That alone is useful..


Conclusion:

The stakes of survey quality are not abstract; they shape whether a study’s claims survive scrutiny, whether subsequent work can build upon them, and ultimately whether the scientific record advances with confidence. Plus, by treating attention checks, exclusion rules, and default platform settings as mere formalities, researchers inadvertently invite noise that obscures true effects and fuels the replication crisis. The metrics outlined — response‑time bimodality, entropy, construct‑relevant variance, and test‑retest reliability — provide concrete, data‑driven signals that can be monitored in real time, allowing investigators to intervene before flawed data become entrenched in the literature That's the part that actually makes a difference..

Equally critical is the structural safeguards that prevent systematic bias from entering the pipeline: pre‑registration of a falsifiable hypothesis, full randomization of presentation order, and the clear separation of manipulation checks from dependent variables. Think about it: the checklist serves as a pragmatic audit trail, ensuring that each of these safeguards has been verified before data collection begins. In practice, when these practices are embedded in the study design, the resulting dataset is far more likely to reflect the phenomenon of interest rather than idiosyncratic participant behavior or platform quirks. Skipping any item is tantamount to compromising the integrity of the entire project, turning what should be a 45‑minute investment in rigor into a costly 450‑participant waste Small thing, real impact..

In sum, high‑quality survey research is attainable when scholars move beyond vague notions of “good data” and adopt a disciplined, transparent workflow anchored by measurable quality indicators and pre‑registered, randomized protocols. The path forward is clear: embed quality controls from the outset, document every step, and make the entire process openly accessible. By publishing these metrics and sharing full study materials, researchers not only protect their own work from irreproducible findings but also elevate the standards of the broader community. Only then can replication become a reliable engine for scientific progress, rather than a recurring lament.

Just Got Posted

New and Noteworthy

Picked for You

Keep the Thread Going

Thank you for reading about Researchers Endeavoring To Conduct An Online Study Should Consider. We hope the information has been useful. Feel free to contact us if you have any questions. See you next time — don't forget to bookmark!
⌂ Back to Home