Our method / Judgment made visible

Every measure
needs a check.

AI helps process research material. Human judgment defines the measure, evaluates errors and sets the limits of interpretation.

Discuss your study

A traceable workflow

From the question
to the final dataset.

The exact design depends on your study. These stages keep the measurement choices visible.

01

Define the construct

Agree what counts as evidence, the unit of analysis and what each code means. For recordings, that includes transcript and timestamp handling. For images and video, it includes observable features and the limits of interpretation.

Output: measurement plan and initial codebook
02

Pilot and settle the rules

Use a development sample to expose ambiguous definitions, source-preparation errors and model failures. Revise the rules here, before evaluating the settled method on a separate validation sample.

Output: documented coding rules and processing workflow
03

Validate against human evidence

Human coders review an agreed sample without seeing model labels where the design requires independence. We report human coding consistency and model-to-human performance separately, using metrics appropriate to the task and sample.

Output: error analysis, disagreements and uncertainty
04

Deliver an inspectable result

Retain the code, prompts, model identifier, settings and versioned outputs. Hand over structured measures with source references, a processing record and limitations. Reruns depend on the available models and approved environment.

Output: analysis-ready data and method record

An example of a coding decision

The important part
is often the disagreement.

A researcher asks why shoppers resist switching to online grocery. This purpose-made response contains two distinct ideas.

“Setting up another account is a hassle. And my usual shop already knows me.”

The first phrase can indicate setup effort. The second may indicate a relationship that switching would disrupt. Both need explicit definitions; neither should be inferred simply because the shopper sounds negative.

Illustrative reasoning only. No client data or fabricated performance statistics.

FOLLOW THE DECISION
Requires checking

One label can miss
a second concept.

An initial coding pass identifies setup effort. It may overlook the existing relationship implied by “my usual shop already knows me.”

Setup effortPresent
Relational switching costNot captured

Select each stage to see how the judgment is documented.

Interpretation has boundaries

Agreement is useful.
It is not the whole story.

A label can agree with a human coder and still fail to represent the intended construct. Validation therefore starts with definitions and source quality, then examines coding consistency and errors.

If labels feed later statistical estimates, we consider whether errors could change the conclusion. An adjusted estimate or sensitivity analysis is included only when the design and validation sample support it.

We do not infer a person’s emotion, intent or identity from appearance or voice alone. Visual and voice material is coded to the agreed, observable research task.

Methods that inform the work

Research behind
the approach.

These publications support task-specific validation and careful use of machine-generated labels.

Start with your research question

What does your
data need to tell you?

Share the question, the material and the deadline.
We’ll work out the right scope together.

Let’s discuss your study