Home
Appearance
Text size
100%
Grid

Renato Valdés-Olmos

Essays

Research borrowed
somebody else’s ladder.

1 September 2026·Operating·~9 min

Go looking for a canonical research career ladder the way you can find Merholz for design or Fournier for engineering or Mehta for product, and you will not find one. There are conference talks and a scatter of company posts. There is no document the field argues about, and that absence is the most interesting fact about the discipline.

It is not that research organizations have no levels. Mine did, at Grammarly and again at Pitch, where research reported into the organization I ran. Researchers were hired at a level, paid at a level, promoted between levels, and the levels mirrored the product design ladder almost exactly: same count, same compensation bands, same expectations about scope and autonomy translated into research language.

That mirroring was the sensible thing to do at the time. Research reported into design, design had the nearest usable standard, and building a fifth document from scratch for a team of that size would have been ceremony. But a borrowed ladder inherits the assumptions of the discipline it was borrowed from, and design's assumptions are not research's.

Every research ladder I have seen is a design ladder with the nouns changed.

What the field has to work from

The whole shelf

This is the whole shelf, and the point of the essay is how short it is. There is no research equivalent of Fournier or Mehta — nothing the field argues about, only documents individual companies happened to publish.

Framework
Year
Shape
What it levels on
2017
Research added to an engineering-shaped matrix
Named competencies at each level
2019
Research inside a shared UX competency set
Skills and performance across UX
2021
Research about careers, not a ladder to run
Observed progression, described
Everything else
Conference talks and company blog posts
Whatever the author’s company happened to use

Two things are true of nearly all of them. Research is carried inside somebody else’s framework rather than given its own, and the ladder stops early: the complaint that there is nowhere left to grow is a structural fact about the documents, not a mood.

I have some sympathy for the borrowing, because I know what the alternative costs. At Honor, where I was the founding designer, the first research program was one I ran myself, in the homes the product was for, before the product had a name. Nobody leveled it, because there was nobody to level it against. When the team grew to twenty across product, brand, design ops and research, the researchers were leveled against the designers, because the designers had a ladder and the alternative was writing one in a week I did not have.

What the borrowed ladder got wrong

Mirroring design was a defensible shortcut and it was also wrong in a specific way. A design ladder measures the progression from executing a screen to owning a surface to shaping a strategy, and each level is anchored in artifacts somebody can look at. Transposed to research, that progression becomes: runs a study, owns a research area, shapes what the company believes.

The first two translate. The third does not, because the thing that separates a strong senior researcher from a strong staff one is not the size of the area they cover. It is whether the organization changed its mind.

At Pitch the research team produced the finding that turned the company’s plan. It was not a large study. It was the right question, asked once, with the nerve to bring back the answer nobody had ordered, and it changed what the company believed about itself. Under the borrowed ladder that work scored as a single study in a single quarter. A researcher can run flawless studies for two years, produce reports everybody praises, and change nothing. Under a borrowed design ladder that person looks promotable: their scope grew, their studies got more complex, their stakeholder list got longer. Under a research ladder that measures the right thing they are stuck, and everyone senior can feel it without being able to point at the row that says so.

The design ladder, with the nouns changed

What happens when a research ladder is transposed from a design one. Four of the rows survive the translation. Two arrive meaning something they cannot measure, and one of them is the row the discipline actually turns on.

The design row
Becomes
Verdict
Scope of work
Studies, then a research area, then a portfolio of areas.
Translates. Territory means about as much in either lane, which is to say less than it used to.
Autonomy
Runs a study with supervision, then without it.
Translates, and runs out just as fast. Three real steps, not six.
Craft depth
Method fluency: interviews, diary studies, surveys, statistics.
Translates, and is the part a model is now fastest at.
Collaboration
Stakeholders served, readouts delivered, teams supported.
Translates, and measures service rather than consequence.
Craft quality
The rigor of the study.
Arrives broken. A flawless study answering the wrong question scores full marks.
Influence
Whether people listened.
Arrives as a popularity measure. The row it should be — whether the company changed its mind — does not exist on a design ladder to be copied from.

The shaded rows are the ones that do not survive the transposition, and they are the two that separate a strong senior researcher from a strong staff one. Everything above them is a measure of activity.

The border is not with design

The other thing the mirror got wrong is where research actually sits. Because research usually reports into design, the assumed adjacency is with design. In practice the live boundary is with data science and analytics, and it has been for years.

The questions are the same questions. Why did this cohort stop coming back. What do people actually do with the feature. Which of these two explanations for the number is true. One side answers with interviews and a synthesis, the other with a query and a chart, and in most companies the two sides sit in different reporting lines and meet at a readout.

What has changed is that the tooling collapsed. A researcher who can write the query is no longer unusual. An analyst who can run six conversations to work out what the drop-off means is no longer unusual either. Synthesis, the thing that was supposed to be the researcher's protected craft, is now the thing most obviously assisted by a machine, which is uncomfortable and worth saying plainly.

It is also, already, a role. When The AI Daily Brief mapped the team of a company running on agents this July, the scout — the one who aggregates signal from the world — was the archetype most often filled by an agent rather than a person. Scouting is most of what a research ladder has been measuring at levels one to three. What it never measured, and what no agent has a stake in, is which question is worth the company’s attention.

Research and Data science

Research keepsStudy design. The question worth asking. The answer nobody ordered.
Now sharedCausal reasoning. Instrumenting a question. Querying the warehouse. Reading a result honestly. Synthesis.
Data science keepsModeling. Inference at scale. What the number can and cannot support.

The live border, and the one no research ladder is drawn against. Synthesis moved into the shared column in about two years, which is why the middle of the ladder is where this hurts.

Research and Product

Research keepsWhether the thing everyone believes is true.
Now sharedExperiment design. Framing what is worth learning. Deciding when the evidence is enough.
Product keepsThe decision the result feeds. What to do when the result is inconvenient.

The border everybody expects research to have with design, it actually has with product — and it is a border about evidence rather than about craft.

What survives

If a model can cluster forty transcripts into themes faster and more consistently than a person, and it can, then thematic synthesis stops being a level. It becomes a step. This is the same collapse that happened to visual execution in design and to estimation in product, and it lands hardest on the middle of the ladder, where most researchers are.

Three things survive it, and they are what I would build a research ladder on.

Question quality. The most consequential act in research is deciding what to ask, and it happens before any method is chosen. A researcher who reliably converts a vague stakeholder anxiety into a question worth twelve interviews is doing something no tool does, because the tool has no stake in the answer.

Evidentiary honesty. Knowing what a finding can carry. Six interviews do not support a roadmap decision and everybody in the room wants them to. The researcher who says so, when saying so kills a plan somebody has already staffed, is operating at a level the borrowed ladder has no row for.

Belief change. Not reports delivered, not studies run, not stakeholders served. Which of the company's beliefs is different because this person worked here, and does it still hold. This is the hardest to measure and it is the only one that separates the top two levels.

The six axes, in research’s language

Three of the six levels. The row a borrowed ladder cannot produce is the third one down.

Axis
R3 · Which problem
R4 · What we refuse
R5 · What outlives us
Decision rights
Decides what is worth studying and what method answers it, without a brief.
Can refuse a study. Can say six interviews will not support this decision, and have that hold when a plan is already staffed.
Sets the evidentiary bar the company makes decisions against.
Judgment density
The room stops arguing about what the data means because they are in it.
Called into decisions with no study attached, for the reasoning rather than the findings.
Teams reason about evidence correctly without them, because the standard is written down.
Opinions held
Brings a point of view about what the evidence means, not only what it says.
Has told the company something it did not want to hear, on the record, and been right about it.
A belief the company holds is theirs, and it still holds.
Range
Writes their own queries. Can read an experiment result and argue with the analysis.
Trusted with product and analytics calls that are not theirs, because they have been right about them.
Routinely in the framing of problems across lanes.
What survives them
A reusable instrument, a panel, a measure the team keeps using.
A standard for what counts as evidence, applied by people who are not researchers.
A belief change that outlasted the team, the reorganization and them.
Ownership
Involved: follows the question past the readout.
Invested: carries an uncomfortable finding to the room that needs it, unprompted.
Invested, and the reference point for what the word means here.

The opinions-held row is the one no borrowed ladder can produce, because it has no ancestor on the design side. It is also the row that makes belief change gradeable rather than aspirational.

Why the field never settled one

I think the reason there is no canonical research ladder is that research teams are small enough to run on borrowed standards and just large enough for that to fail quietly. At six researchers, the design ladder with the nouns changed works well enough. At twelve it does not, but the failure shows up as an unexplainable promotion decision two years later rather than as a visible gap, so nobody traces it back to the borrowed document.

The other reason is that writing one down means committing to belief change as the measure, and that measure is uncomfortable for exactly the people who would have to write it. It makes a researcher's level partly dependent on whether the organization listened, which feels unfair, and is unfair, and is also true.

The fix is not to score them on something they fully control. It is to hold the organization to the same row: if research is not changing what the company believes, that is a finding about the company, and it should appear in somebody else's review.

The whole system

You can read the argument for free. Running it is the hard part.

An essay can tell you the six rows stopped working. It cannot sit in the room in November when two managers disagree about the same person and neither can say why. That is what the workbook is for: one shared core across Engineering, Product, Research and Design (EP(R)D), six levels defined by what a person decides, and the mechanics to grade against them without the cycle turning into a negotiation.

Get the workbook · $39 →

Decides
Design
Research
Product
Engineering
1How
D1
R1
P1
E1
2What
D2
R2
P2
E2
3Which problem
D3
R3
P3
E3
4What we refuse
D4
R4
P4
E4
5What outlives us
D5
R5
P5
E5
6What we will not do
D6
R6
P6
E6

Six levels defined once, by what a person decides alone. The shaded band is the overlap, where most of the work now belongs to no single lane. The marks to the left of each cell are doors: a move sideways at the same level, which is a transfer and not a demotion.

Written to be opened during a cycle rather than read once: on the page, as a PDF set to print, and with the templates as a spreadsheet.

Define it

01Why the old ladders stopped workingSample chapter
02The spine: six levels, six axes
03Four lanes, and what each one keeps
04Overlap, doors and the management fork

Grade against it

05Behaviors: the full matrix
06Craft: what to score, at what depth
07Evidence: the decision trail
08Grading, with worked examples
09Nine ways grading goes wrong

Run the cycle

10Running a cycle
11The calibration room
12Promotion cases
13Templates
Three people, graded in fullA designer who should not be promoted yet. An engineer who looks excellent and is not a level 5. A researcher moving into product. Marked axis by axis, with the argument written out the way you would have to make it in the room.
Five templates you can use in NovemberEvidence log, self-assessment, calibration scorecard, promotion case, transfer record. Short enough that people actually fill them in.
The rule that stops grade inflationDecision rights carries a veto. Below on that axis and the level is not held, whatever the other five say. Most frameworks leave this implicit, which is how everybody ends up a four.

v1. The cycle, the calibration room and the evidence standard are the ones I have run at companies of 25, 250 and 2500 people; the six axes are new. One payment, every revision by email.

Compensation$39For CFOs and Finance teams setting pay for Engineering, Product, Research and Design, with EPD leaders supplying the role and market evidence behind every decision.
Development$39For Engineering, Product, Research and Design leaders and People teams to run together: turn review evidence into funded growth work, protected time and explicit manager commitments.
The new standard$117$99A shared standard for levels, pay and growth. All three workbooks, their printable documents, 24 editable templates and reference sheets, plus three worked examples.

One of four, each taking a function through the same framework; the system they share is set out in full in the workbook. The others are on design, engineering and product. Written from having inherited, run and rebuilt leveling and review cycles at three companies of very different sizes, and against the public frameworks that preceded them: the levels framework in Org Design for Design Orgs by Peter Merholz and Kristin Skinner, Rent the Runway’s engineering ladder, and Ravi Mehta’s product competency toolkit. Figures from the AI in Design Report 2026 and Figma’s State of the Designer 2026.

© 2026 Renato Valdés-Olmos