Schedule a demo
Resources / Guide

How to choose eye tracking glasses for research

Wearable eye trackers all produce a gaze point. The differences that matter show up later, when you try to turn that gaze point into a result somebody will accept.

Last updated 18 September 2026

Summary

Seven questions decide the choice. Does the participant have to move? What are you measuring, and how fast does it move? Do you need to know the object, or is a point on a video enough? Can the data leave the building? Will participants be wearing their own glasses? What happens when a reviewer challenges a number? And how much setup time can you afford per participant?

Answer those seven honestly and the shortlist usually writes itself. This guide goes through each one, including the cases where the answer points away from the kind of system we build.

1. Does the participant have to move?

This is the category question and it is worth being strict about.

If your participant sits in front of a monitor and the stimulus is on that monitor, a screen-based tracker is usually the better instrument. Higher sampling rates are available, there is no head geometry to model, the analysis is simpler, and it costs less. Buying wearable glasses for a screen study is paying for mobility you will not use.

Wearable glasses exist because some tasks cannot be brought to a desk. Driving. Flying. Surgery. Walking a shop floor. Operating a machine. Training on a real piece of equipment. In those studies the participant moves, the scene moves, and there is no screen to reference.

One thing worth knowing before you decide: some wearable systems can also map gaze onto a stationary screen, which lets one instrument cover both kinds of study. If your lab runs a mix, it is worth asking about rather than buying two systems.

2. What are you measuring, and how fast does it move?

Sampling rate is not a specification to compare for its own sake. It follows from what you are measuring.

SLOW

Fixations and dwell time

Where attention landed and for how long. A fixation typically lasts 200 to 300 milliseconds, so 30 or 60 Hz gives you enough samples to find it. This is most applied research, and a lower rate is genuinely fine.

FAST

Saccades

A saccade lasts roughly 30 to 80 milliseconds. At 60 Hz you get two to five samples across the whole movement, which is not enough to characterize its shape, peak velocity or duration with confidence. At 180 Hz you get six to fifteen. If saccade dynamics are your dependent variable, sampling rate decides whether the study is possible.

3D

Depth and vergence

This needs both eyes measured independently. Some systems track one eye and estimate the other, or track both but can only report a gaze pixel position on a video image. Some infer depth from a scene depth camera, which reports how far away an object is rather than where the two eyes converged. None of those gives you real vergence. If depth matters, ask directly whether each eye is measured, with eye position and gaze direction, in three-dimensional space.

3. Do you need the object, or is a point on a video enough?

This is the question most buying guides skip, and it is usually the one that determines how much work the study becomes after the data is collected.

Every wearable tracker gives you a gaze point on a scene camera video. That tells you where on the video image the participant looked. It does not tell you what they looked at. Turning one into the other is the real cost of a wearable study, and there are five ways to do it.

Manual coding

A person watches the video and writes down what was being looked at, frame by frame. It works, it is the fallback for anything unusual, and it is expensive. Budget hours of coding per minute of recording, and expect to run reliability checks between coders.

Static areas of interest

You draw regions once and apply them. Fine when the scene does not move relative to the camera, which on a head-mounted camera means almost never.

Automated moving areas of interest

The software tracks a region you defined as it moves through the video. This is a large improvement and it covers a lot of applied research. Its limit is that it is still working in the camera image, so an object that leaves frame, gets occluded, or changes appearance can break the track.

Live object recognition

An AI model identifies objects in the scene and the system reports which one gaze landed on, as it happens. The useful version of this is one you can train on your own objects, because a generic AI model will not know what a specific instrument panel or a specific product is. Ask two questions: can I train it on my objects, and does it run during the session or only afterwards?

Gaze in room coordinates

The most complete answer. Head position and orientation come from an external tracking system. Combined with binocular eye data, gaze becomes a point in the room rather than a pixel position on a video frame, and it can be reported against real surfaces and objects whose positions you have defined. The same coordinate frame applies to every participant, which means you can aggregate across a study without re-coding anything.

The cost is setup. You need a motion capture or head tracking system, and you need to define the surfaces. How gaze is placed in three dimensions walks through how that is built. This is the right answer for a cockpit, a vehicle mockup, a control room or a simulator, where the geometry is fixed and the question is which instrument was checked and when. For somebody walking around a supermarket, live object recognition such as Argus AI LMOT or LAOI that names the object as the glance happens, is the better instrument for that study.

4. Can the data leave the building?

Ask this before you build the shortlist, not after you have chosen.

Eye tracking data is unusually sensitive. A scene camera records whatever the participant was looking at, which in a hospital, a factory, a defense facility or a school means faces, screens, documents and equipment that were never consented to appear in a recording.

Systems differ more than people expect. Some do gaze computation on your machine. Some upload recordings to a vendor service for processing. Some work offline but need to reach a license server periodically. Some store recordings in a cloud account by default and offer local storage as an option.

None of those designs is wrong. They are different products for different situations. But you need to know which one you are buying before somebody in your institution's information security office asks you, because at that point the answer determines whether the study happens.

Five questions to put to any vendor in writing:

If you work anywhere with an export control regime, a classification, a hospital information governance process or a corporate data residency policy, put those questions in the procurement documents. For how one system answers them, see ETVision.

5. Will participants be wearing their own glasses?

A large share of adults wear some form of corrective eyewear, and the proportion rises steeply with age. In a driving study, an ageing study, a clinical study or anything involving a working professional population, this is not an edge case. It is most of your sample.

There are three ways to handle it, and they are not equivalent.

INSERTS

Corrective lens inserts

The tracker takes prescription lenses fitted to the frame. This works, and it means collecting prescriptions in advance, ordering and managing a set of lenses, and accepting that astigmatism and progressive prescriptions are handled poorly or not at all.

EXCLUDE

Exclude glasses wearers

Common, rarely stated plainly in the methods section, and it silently biases the sample. If the population you care about mostly wears glasses, you have replaced your research question with a different one.

OVER

A frame worn over their own glasses

The participant keeps their own correction, including astigmatism and progressives. Nothing has to be ordered in advance and nobody is excluded. The trade is bulk, and the eye camera has to be positioned to see past the participant's lens, which means the optical design has to allow for it rather than treating it as an afterthought.

Whichever route you take, decide it before recruitment, and write it into the protocol. Finding out on the first session that a third of your participants cannot be tracked is an expensive way to learn this.

6. What happens when somebody challenges a result?

At some point a reviewer, a supervisor, a regulator or an opposing expert will ask how you know a particular fixation was on the thing you say it was on. The answer cannot be "the software said so".

What you need, in order of how often it matters:

7. How much setup time can you afford per participant?

This is the section where the field's marketing is least clear, so it is worth going slowly.

The short version, for a system like ETVision: auto calibration takes about five seconds. The rest of this section is about why that number matters less than what those five seconds buy you.

What calibration actually does

Your visual axis, the line from what you are looking at to the fovea, does not coincide with the optical axis of your eye. The angle between them is person-specific and commonly a few degrees. Corneal curvature varies between people too. A calibration measures those offsets for the individual sitting in front of you. That is the whole purpose. It is not a ritual.

What "calibration-free" means

Instead of measuring your eye, the system applies a general model, typically learned from a large population of eyes. For an eye close to the population average, this works well. For an eye that is not, the error is larger, and the system has no way to tell you which one it just measured. That is not a defect. It is the direct consequence of using an average instead of a measurement.

When calibration-free is genuinely the right choice

There are real cases, and they are not rare:

In those studies, a system that produces usable data with no procedure is worth more than a system that produces better data you were never going to be able to collect.

In practice

Most researchers running studies where the accuracy number matters perform a calibration anyway, including on systems sold as calibration-free, because it is the only way to know what the accuracy was for that participant. "Calibration-free" honestly means "will produce data without calibration". It does not mean "equally accurate without calibration", and the two get conflated in almost every comparison you will read.

Also worth knowing: the realistic alternative is not a fifteen-minute procedure. Auto calibration takes about five seconds. The gap between no calibration and auto calibration is much smaller than the way it gets discussed suggests, and it buys you a number you can defend.

So ask it this way: not "does this need calibration", but "what is the accuracy for this particular participant, and how would I know?"

8. Putting it together

The participant sits at a screenA screen-based tracker, or glasses can also do screen studies
The participant moves through a real environmentWearable glasses
You are measuring saccade dynamics or pupil responseHigh sampling rate, and pupil diameter measured from an ellipse fit
You need depth or vergenceBoth eyes measured independently, with eye positions and gaze directions, in three-dimensional space
You need the object, not where on the frameAutomated AOI tracking at minimum, object recognition or room coordinates at scale
Fixed geometry: cockpit, console, mockupGaze in room coordinates with head tracking
Data cannot leave the siteLocal processing, offline operation, no license server
Your population wears glassesA frame worn over their own eyewear
The result will be reviewed or challengedRaw eye video, frame replay, post-hoc recalibration, deterministic output, not a black box
You cannot calibrate your participantsA calibration-free system, accepting the accuracy trade

9. Frequently asked questions

What sampling rate do I actually need for eye tracking?

For fixations and dwell, 30 to 60 Hz is usually enough. For saccade dynamics, or smooth pursuit, you want well above 120 Hz. A saccade lasts 30 to 80 milliseconds, so at 60 Hz you are describing the whole movement with two to five samples.

Is 0.5 degrees better than 1 degree of eye tracking accuracy?

Only if both numbers were measured the same way. Accuracy figures are not standardized across the field. Ask what was measured, over what field of view, with how many participants, and how error was defined. A number without a method behind it is not comparable to anything.

How long does eye tracker calibration take?

On ETVision, auto calibration takes about five seconds with automatic feature detection. That number is worth holding on to when a system is marketed as calibration-free, because the real choice is not five seconds against a fifteen-minute procedure. It is five seconds against an accuracy figure you cannot check for the participant in front of you.

Can one wearable system replace a screen-based tracker?

Sometimes. Some wearables, such as Argus ETVision, can map gaze onto a stationary screen and analyze it against stimuli, which covers a lot of screen-based work.

Do I need motion capture for 3D gaze in the real world?

For gaze in real room coordinates, yes. Head position and orientation have to come from somewhere. There are lighter-weight head trackers as well as full motion capture arrays, and the right one depends on the size of the space and the precision you need.

How do I handle participants who wear glasses?

Decide before recruitment. Either fit corrective inserts, exclude them and state it in your methods, or use a frame designed to be worn over their own glasses. The third option is the only one that does not change your sample.

What should I ask an eye tracking vendor about data handling?

Where gaze is computed, whether the software runs fully offline, where recordings are stored by default, whether a license server must be reachable, and what leaves the machine. Get the answers in writing before procurement, not after.

Related

GLASSES

ETVision

The wearable system underneath it. Both eyes measured independently at 180 Hz, which is what makes vergence a measurement rather than an estimate.

TECH

3D Gaze in the Real World

How gaze is placed on real objects in a room, and what a world model has to contain for that to work.

FAQ

Frequently asked questions

Prescription glasses, cloud processing, accuracy, motion capture, and support for legacy ASL systems.

Working through these questions for a study?

Describe it and we will tell you what we would use, including when the answer is not us.

Contact us