GMU: Can assistive technology compensate for poor web accessibility?
Capstone Project at the Visual Attention and Cognition Lab, Department of Psychology — part of a 2-year M.A. in Human Factors and Applied Cognition.
A capstone research project at George Mason University's Visual Attention and Cognition Lab, using Tobii eye-tracking to prove that WCAG principles are essential — not optional — for creating websites usable by everyone. When 97.4% of top websites fail accessibility guidelines, the burden falls on assistive technologies that were never designed to compensate.
Overview
Capstone Project at the Visual Attention and Cognition Lab, Department of Psychology — part of a 2-year M.A. in Human Factors and Applied Cognition.
If 97.4% of websites fail accessibility, can assistive technology bridge the gap?
In 2021, WebAIM analyzed the top 1,000,000 most visited websites and found that 97.4% of homepages contained features that don't conform to the Web Content Accessibility Guidelines (WCAG). This places a significant burden on assistive technologies — tools primarily designed to work with accessible websites, not compensate for inaccessible ones.
This eye-tracking study was the capstone of my two-year M.A. in Human Factors and Applied Cognition at George Mason University, conducted in the Psychology department's Visual Attention and Cognition Lab. Eye-tracking was the right instrument for the question: it surfaces where attention actually goes, moment to moment, in a way self-reported surveys and interviews can't.
The study asked two research questions:
RQ1 — If a website adheres to WCAG accessibility guidelines, is costly assistive technology (such as magnification software) still necessary?
RQ2 — If a website doesn't adhere to WCAG guidelines, can assistive technology compensate for the lack of accessibility?
The short answer to both turned out to be no: participants navigating an accessible website didn't need assistive technology at all, and on an inaccessible website, assistive technology couldn't rescue them. The evidence for that claim — behavioral, attitudinal, and gaze data — is below.

- 97.4%Of top 1M website homepages fail WCAG accessibility
- 3×U.S. adults with disability less likely to use the internet
- 6 minSaved on accessible vs. inaccessible tasks (n=3 pilot, directional)
- 88.33SUS score for accessible website (industry avg: 68; n=3 pilot)
Why It Matters
Vision disability is the fifth most common disability in the U.S.
Severe vision loss presents problems with reading, navigating, using assistive technology, and using the internet. Low vision is most commonly caused by macular degeneration, cataracts, diabetic retinopathy, and glaucoma. U.S. adults with a disability are 3× less likely to use the internet (Pew Research Center, 2017), are over 3× less likely to be employed (BLS, 2022), and report higher levels of social isolation compared to their abled counterparts.

The Assistive-Technology Gap
The industry's implicit answer to inaccessible websites is "let assistive technology handle it." But most assistive technologies are expensive and hard to use for browsing, which compounds low-vision users' barriers to employment and social integration. Screen magnifiers — among the most common assistive tools — have three well-documented problems:
- They impede screens globally
- They disrupt spatial orientation
- They lead to excessive scrolling
If assistive technology can't actually compensate for inaccessible design, then WCAG compliance isn't a nice-to-have — it's the only mechanism that works. That's the hypothesis this study put to the test.
Study Design
A within-subjects experiment: one variable, four measures, real constraints.
The independent variable was WCAG compliance: each participant completed the same information-finding tasks on an accessible website and an inaccessible one. The dependent measures were task completion time, task completion rate, System Usability Scale (SUS) scores, and gaze behavior from the eye tracker — two behavioral measures, one attitudinal, one attentional, so any conclusion would have to survive convergent evidence rather than lean on a single metric.
I employed a within-subjects design so each participant served as their own control, cancelling out individual differences in reading speed and web fluency — the highest-variance factors in a small sample. To mitigate order effects, condition order was counterbalanced: two participants started on the accessible website, one on the inaccessible one.
The design also had to absorb real-world constraints: a limited budget ruled out recruiting low-vision participants, the lab schedule allowed exactly one week for data collection, and the Tobii eye-tracking rig had to be set up from scratch with no prior lab infrastructure. Those constraints shaped the method rather than the question — simulation instead of clinical recruitment, three participants instead of a full panel, and a topic scoped down to the two most defensible research questions.
Simulating Macular Degeneration
I selected the most common cause of low vision — macular degeneration — and simulated its central symptoms, a central scotoma with blurred vision, using the Silktide Disability Simulator Chrome extension. The scotoma tracked the participant's mouse position, occluding the point of regard the way a real central blind spot follows the fovea.
Participants
Three graduate students with normal or corrected-to-normal vision took part in the study.
Instruments
- Tobii eye tracker — recording fixations and saccades on both websites to capture attention allocation directly
- Silktide Chrome extension — simulating a central scotoma and blurred vision attached to mouse movements
- Zoom magnification software — docked at the top of the screen to prevent occlusion, magnifying text at the mouse's location

The Two Websites
Identical information, opposite WCAG compliance.
The accessible condition used the New South Wales Government website — built top to bottom on WCAG principles. For the inaccessible condition, we rebuilt the same content as a mockup website with those principles deliberately stripped out. Same information architecture, same tasks — the only thing that changed was the design's compliance.
1. Typography & Contrast
The accessible site used sufficient contrast and readable font sizes. The inaccessible version used insufficient contrast, cursive display fonts, and all-caps text — individually small decisions that compound under a scotoma.


2. Dropdown Menus
The accessible site used multicolumn, multirow dropdowns with hover-state color changes and highlighted tab headers. The inaccessible site used single-column dropdowns with no visual feedback on hover or selection — participants' gaze data later showed exactly what that difference costs.


3. Column Layout
The accessible site used even, single-column layouts that read cleanly under a magnification lens. The inaccessible site used uneven, unnecessary columns — under magnification, a reader loses the line they were on every time the column breaks.
In the Lab
Think Aloud protocol with eye-tracking — observing the struggle directly.
Participants arrived at the laboratory and verbally agreed to be recorded. Each calibrated their pupils with the Tobii eye tracker, which then compiled screen recordings of fixation locations across both websites. Participants were shown the Silktide simulator (scotoma attached to mouse movements) and the Zoom magnifier docked at the top of the screen, and were instructed to navigate and read as they normally would, thinking aloud as they worked.
Tasks
- Find a place to get swabbed for COVID-19
- Look up regulations about COVID-19 in the area
- Find support for people with a disability in COVID-19
I took structured observation notes as participants worked, recording their verbalizations, reactions, and behaviors for later coding. In the gaze plot below, each red dot is a fixation — a point where the eye stops to process information, sized by dwell time — and the connecting lines are saccades, the rapid jumps between fixations.

Quantitative Findings
Every behavioral and attitudinal measure favored the accessible website.
Participants averaged 6 minutes faster on the accessible website across the three information-finding tasks — 2:06 total versus 8:17, roughly a 4× difference.

Task Completion Rates
- 3/3 participants completed all 3 tasks on the accessible website
- 2/3 participants failed to complete at least one task on the inaccessible website
System Usability Scale (SUS)
- Accessible website: 88.33 — well above the industry average of 68 (usability.gov)
- Inaccessible website: 64.17 — below the industry average

On Statistical Significance
With n = 3, a paired t-test on task times did not reach significance at α = .05 — and reporting anything else would be malpractice. What a three-person pilot can establish is directional evidence: every participant was faster on the accessible site, every attitudinal score pointed the same way, and the gaze recordings show why in mechanistic detail. Convergence across behavioral, attitudinal, and attentional measures is what makes pilot data decision-grade even when inferential statistics can't yet be.
Qualitative Findings
What participants said — and what their eyes revealed.
We coded the session notes and Think Aloud audio against the eye-tracking recordings using the HyperResearch qualitative analysis tool, tagging recurring behaviors and reactions across both website conditions.

Session notes and Think Aloud audio coded in HyperResearch; text size reflects how often each code recurred across participants.
On the Accessible Website
Participants reported finding the site "pretty easy" and "clearer." They navigated without needing the magnification reader at all, using their peripheral vision to read around the simulated scotoma. The dropdown "isn't flying around everywhere" — the multicolumn layout allowed scanning without losing context.
On the Inaccessible Website
Participants struggled significantly. "I'm using the shadow to see." When they tried the magnifier, it did not help — "I'm not sure what this is so I would definitely need the magnification for this part. Which really doesn't help considering that it's still really small." The dropdown "was the same color as the background," making navigation nearly impossible.
The eye-tracking data showed larger, more erratic fixation patterns on the inaccessible site — participants' eyes were searching, not reading. The bigger the fixation dot, the longer a participant stared at something they couldn't parse.
Answers
Assistive technology is unnecessary for accessible websites — and cannot rescue inaccessible ones.
RQ1 — answered no. On the WCAG-compliant website, every participant completed every task without touching the magnification software. Peripheral vision around the scotoma was enough, because the design gave it enough contrast, size, and structure to work with.
RQ2 — answered no. On the inaccessible website, the same participants reached for the magnifier and it didn't save them: magnifying an undersized, low-contrast element produces a bigger unreadable element. Accessibility failures live in the design, not in the tooling around it.
Three design implications fall directly out of the data:
1. Contrast is the primary barrier. WCAG requires a minimum 4.5:1 contrast ratio for normal text — a bar the inaccessible site fell well short of. Contrast directly drove the reading difficulty behind the 6-minute gap, and it's one of the cheapest accessibility issues to fix.
2. Dropdown format matters for low vision. Scotomas make single-column dropdowns nearly unusable; multicolumn, multirow formats with hover feedback are substantially more navigable.
3. Small graphics and fonts fail even with magnification. No magnification reader can clarify an undersized element — the fix belongs to the designer, not the user's toolkit.
For practitioners, the conclusion is blunt: WCAG compliance is a baseline, not an enhancement. Assistive technology assists — it does not repair.
Limitations & Next Steps
A rigorous study names its own weaknesses.
Sample
Three participants — all graduate students in their twenties, all comfortable with technology — is a pilot, not a panel. Younger, tech-fluent users likely understate the difficulty a representative low-vision population would face, which means the observed gaps are more plausibly a floor than a ceiling. Nielsen's guidance suggests 5+ users for qualitative testing; inferential claims would need substantially more.
Simulation vs. Reality
A simulated scotoma approximates the optics of macular degeneration but not the lived adaptation — real low-vision users have years of compensatory strategies. The natural next step is replication with visually disabled participants across age groups.
Instrumentation
The SUS measures general usability, not accessibility specifically, and its validity is strongest at 12+ respondents — a dedicated accessibility instrument would fit a follow-up better. On the hardware side, an EyeLink tracker with Weblink software would add heat-map generation and area-of-interest analysis that the Tobii setup couldn't produce.
Scope Under Constraint
A one-week testing window left no room to explore a topic this broad — only to narrow it fast. I cut scope down to the two most defensible research questions and moved straight into execution, adjusting the plan as constraints surfaced rather than following one fixed from day one. That editing instinct — protecting the validity of a small study instead of overreaching — was the most transferable lesson of the capstone.
Usabilathon 2022
Second place at a HelloFresh-sponsored UX marathon — research under a clock.
The capstone's methods weren't confined to the lab. In November 2022, I competed in the Usabilathon 2022, a UX design marathon sponsored by HelloFresh, and placed second — on a team of five (a project manager, a UX researcher, and three UX designers, with a HelloFresh mentor advising). I worked across both UX design and research tracks: emphasizing, ideating, building wireframes and prototypes, and conducting the user testing.
The brief: HelloFresh's in-app "Market" — the add-on grocery store inside its meal-kit app — was buried five navigation levels deep, and users simply never found it. Working mobile-first, we ran a compressed version of a full research cycle:
- Subject-matter-expert interview with HelloFresh's UX researcher to anchor business goals before touching pixels
- Cognitive walkthrough of the existing flow, screenshotting every step to isolate pain points
- Competitor IA analysis (Home Chef, Sunbasket) that motivated raising Market's hierarchy level
- Task analysis — re-mapping the user flow before and after the proposed structure
- Two rounds of usability testing: a six-participant Think Aloud session on the Figma prototype during the event, then a post-hackathon comparison round against the live app
The redesign moved Market onto the homepage, made Meals vs. Market distinguishable, and added search to the meal-selection flow. In the follow-up evaluation (N = 11), the prototype cut time on task by 64.5% and user error rates by 27.4%, while raising add-on conversion in-session by 18.2% — reported, as with the capstone, with the honest caveat that a sample that size doesn't reach statistical significance at α = .05.


Read the full HelloFresh case study or flip through the presentation deck.
Graduate Research
One study in a broader research practice — two years of Human Factors and Applied Cognition.
The capstone was the endpoint of a two-year M.A. in Human Factors and Applied Cognition in GMU's Psychology department — a program built on applying cognitive science to real systems, from human-computer interaction and cognitive engineering through a heavy quantitative methods core. The coursework produced its own line of research along the way: each study below leaned on a different layer of that training, and together they're why the capstone's method choices — instruments, designs, honest statistics — weren't improvised.
Smart-Home Apps: A Full Experimental Design
For research methods training, I designed a usability study of smart-home apps — Amazon Alexa vs. Google Nest — as a complete experimental proposal: 60 participants randomly assigned between apps, five representative performance tasks (from device onboarding through multi-room routine automation), and validated usability instruments, grounded in Nielsen's heuristics and the smart-home HCI literature.
Emotion in Interface Design: An Appraisal-Theory Research Proposal
Usability metrics stop at whether people can use an interface; a second proposal asked what interfaces make people feel. Building on Roseman's appraisal theory of emotions, I designed an empirical study of emotional experience on shopping websites: a research model linking four interface features — information structure, navigation and orientation, text, and visual layout — to three cognitive appraisals and six shopper emotions, from joy and liking to frustration and fear. The method ran in two phases: structured interviews with 50 psychology students to screen which of Roseman's 17 emotions actually occur while shopping on Amazon, then a 100-participant field survey with measurement items adapted from IBM's web design guidelines on 9-point Likert scales. The analysis plan — structural equation modeling in Mplus, with convergent and discriminant validity checked against Fornell and Larcker's criteria — treated the pipeline as part of the design, not an afterthought.
Designing for Well-being: From Usability to Flourishing
In PSYC 768 — a seminar on human-systems and human-robot interaction built around nomological networks and conference-style research writing — I took the question one step further: past usability, past emotion, to whether technology leaves its users better off. The resulting paper, Designing for Well-being in Digital Experience: A PERMA Perspective, translated Seligman's five-element well-being model (Positive Emotions, Engagement, Relationships, Meaning, Accomplishment) into concrete design and evaluation criteria for digital products, weighing it against Self-Determination-Theory frameworks like METUX and arguing for evaluation as a cycle of assessment and adjustment rather than a one-time audit.
Quantitative Training
The statistics under all of this ran through the program's methods core. PSYC 652 — Analysis of Variance — covered between-subjects, within-subjects, mixed, and random-factor designs, computed both by hand and in SPSS from Keppel and Wickens' Design and Analysis, ending in a full write-up of an ANOVA study on a real dataset. That's the training behind the capstone's design decisions: knowing why a within-subjects design cancels individual differences, why order needs counterbalancing, and why n = 3 means reporting direction rather than significance. HyperResearch rounded out the toolkit on the qualitative side — systematic code-and-retrieve analysis of Think Aloud audio and observation notes, so the qualitative claims were as auditable as the quantitative ones. The through-line of the whole program: pick the instrument after the question, and report the uncertainty with the result.
Go Deeper
Read the full capstone research report for the complete methodology, literature review, and data behind this study — or the eye-tracking and low vision presentation for the same findings deck-style. The study's boards and supporting artifacts are collected in the Figma file. The coursework research is available in full: the smart-home usability study, the shopping-website emotion research proposal, and the PERMA well-being paper.



