This website uses cookies

Read our Privacy policy and Terms of use for more information.

The views expressed here are solely my own. They do not represent the opinions, positions, or policies of any current or former employer, client, or affiliated organization.

I recently wrote about a practical exercise where I underperformed, despite having more than a decade of real experience solving the exact problem I was asked about. That article was about turning that stumble into a framework. This one is about the stumble itself, because I think the format of that exercise deserves scrutiny, not just my reaction to it.

Here's my claim: simulated, individual, timed evaluations are often a poor proxy for how someone actually performs, though not always, and it's worth being honest about when the discomfort is a design flaw versus simply the cost of clearing a necessary bar. I have one case that clearly shows the flaw, one that's more ambiguous, one where a human judgment call overrode the test and turned out to be right, and one counterexample, 11 years old, that proves the format itself doesn't have to be this way.

I'm not fully certain I'm right about this. But the pattern has repeated enough times, across different companies and formats, that I think it's worth naming plainly.

Why Simulated Environments Distort Performance

There are five reasons I keep running into, and I think they compound each other rather than acting independently.

Low ecological validity. The brain registers when a situation isn't "real," and it processes information differently as a result. Some people need an authentic context, real stakes, real consequences, to activate their abilities at full capacity. Strip that away, and you're not measuring the skill anymore; you're measuring how well someone performs in the absence of the conditions that skill normally depends on.

Evaluation anxiety. In an interview or a practical test, you know you're being watched and judged. That awareness consumes working memory and reduces performance, even when you know exactly how to do the task. This isn't a confidence problem. It's a cognitive load problem, and it affects people who are otherwise highly competent just as much as anyone else. It's the same mechanism behind a phrase almost everyone has heard at some point: "just act normal." The moment someone says that to you, normal is exactly what you stop being able to do, because the instruction itself makes you aware that you're being observed, and that awareness is precisely what breaks the thing you were doing effortlessly a second ago.

Absence of context. In real work, you have documentation, colleagues, time to research, internet access, prior emails, institutional memory. A simulation strips away most of that, even though in real life those resources would exist and would be part of how you actually solve the problem. Removing them doesn't make the test "purer", it makes it test something different: your ability to perform without support, which is rarely the job itself.

Objective shift. At work, the objective is to solve a problem. In an interview, the objective quietly becomes "demonstrate that I know how to solve a problem." That shift can affect people very differently, some perform better under a demonstration frame, but for others, it introduces a layer of self-monitoring that actively gets in the way of the thinking itself.

Natural vs. artificial processing. Some people do their best work when they can explore a problem, ask questions, test hypotheses, and iterate. A timed practical test usually expects a fast, linear answer. If your strength is depth through iteration rather than speed through certainty, the format itself works against you before you've said a single word.

A Second Data Point: The Canadian Road Test

This pattern isn't limited to job interviews. I ran into it again, in a completely different domain, when I moved to Canada and had to take the driving road test.

I'd been driving in Colombia for about 20 years by then, no accidents, no infractions. Driving on Colombian streets is, frankly, a much higher-difficulty environment than driving in Canada: denser traffic, far less predictable behavior from other drivers, infrastructure that forces constant improvisation. By comparison, driving in Canada is rookie level.

I failed the road test twice before passing it on the third try.

Not because I couldn't drive. A standardized test, evaluated by a stranger sitting next to me with a clipboard, checking a fixed sequence of maneuvers under artificial scrutiny, doesn't capture 20 years of real-world driving competence, it captures something narrower.

And to be fair to the format itself: I'm not sure there's a better approach available for this specific case. Rules of the road are rules of the road, and a controlled test is probably the most reasonable way to verify that someone actually knows them, regardless of how many years they've been driving somewhere else. Unlike the interview scenario, this isn't really a case of bad design, it's closer to a case where the format is doing exactly what it's supposed to do, and the discomfort is simply the cost of clearing that bar. You don't get to negotiate with it; you study, you practice, and you pass it, whether it feels representative of your real skill or not.

When Someone Overrides the Test With a Hunch

There's a more recent case that, in some ways, is the strongest evidence I have, because it's still unfolding.

I failed the final interview for my current job. The questions in that round hit exactly the kind of wall I described above, the format exploded my thinking in real time, and I knew it wasn't going well while it was happening. By every formal measure of that specific evaluation, I should not have gotten the offer.

I got it anyway. I'd had a second interview earlier in the process with the person who would become my manager, and after the rejection came through, I was told she'd liked my profile enough that she pushed for the hire regardless. I attribute that to a hunch on her part, a read on fit and capability that didn't come from the formal test, but from a different kind of signal entirely.

Almost 18 months later, I can say, without trying to sound self-congratulatory, but also without downplaying it into false modesty, that my performance has been strong. In some areas I've gone beyond what was expected of the role, and I've been given more responsibility as a result.

I bring this up not to argue that interviews don't matter, but to point at something specific: the final formal test said no, and a human judgment call said yes, and 18 months of real evidence has sided with the hunch, not with the test. That's not proof that hunches always beat structured evaluation. It's one data point. But it's a very clean one, because unlike the group exercise I'll describe next, or the road test above, this one has an ongoing, measurable track record attached to it, not just a memory of how a moment felt at the time.

The Best Selection Process I've Ever Been Through

More than 11 years ago, I went through a selection process at a Fortune 500 consumer goods company that I still think about as the best-designed evaluation I've experienced, and it's the clearest counterexample to everything above.

The setup: 13 candidates in a room. We were given a fictional problem, not related to the actual job, not testing domain expertise, just a scenario with clear context and clear rules. Each person had to propose a solution and then defend it against the others in open discussion. Four evaluators sat in the room watching.

What made it work wasn't the format on paper, group exercises can be just as artificial as individual ones. What made it work is that at some point, we forgot we were being evaluated.

Because the problem had nothing to do with our actual jobs, there was no "correct professional answer" to perform. There was just a scenario to reason through, in real time, in front of peers who were doing the same thing. The objective never shifted to "prove I'm good at this", it stayed on "actually solve this and convince these people I'm right." That's precisely the trap the individual, job-related simulation falls into, and precisely what this format avoided.

Three of us were selected out of the 13. And here's the part that matters most to me: I ended up working closely with those two other people for almost seven years afterward. Looking back, it was a genuinely good call, for the team, and for the company. One data point doesn't prove a method works universally. But it's a strong signal that this specific format captured something real about how each of us thought, argued, and negotiated, something that an individual timed test, in my experience, rarely does.

What Actually Makes an Evaluation Format Work

Comparing that experience against every timed individual simulation I've done since, a pattern emerges. The formats that work well seem to share a few traits:

  • They displace the object of judgment. When the scenario isn't your actual domain, you're not defending your competence, you're just thinking out loud. That alone dissolves a huge share of evaluation anxiety.

  • They allow real interaction, not a monologue against a clock. Defending an idea against pushback from peers reveals reasoning, flexibility, and negotiation skill in a way a solo presentation never can.

  • They give you enough context to actually think, not just enough to perform. Clear rules and a well-defined scenario aren't "cheating", they're what lets the exercise measure thinking instead of measuring stress tolerance.

  • The stakes feel low enough to forget you're being watched, and high enough to actually try. That balance is hard to design, but it's the whole point.

What This Means for How Companies Evaluate People

I want to be careful here: this isn't an argument that I should be excused from ever being tested. It's an argument that the format of the test determines what you actually measure, and a lot of companies default to individual timed simulations because they're easy to standardize and score, not because they're the best predictor of real-world performance.

If a company wants to know how someone thinks, argues, and collaborates under real conditions, a fictional group scenario with clear rules and real interaction will very likely surface more truth than a solo exercise where the candidate is quietly trying to guess what "correct" looks like to an evaluator.

I don't think every hiring process needs to look like that room from eleven years ago. But I do think it's worth asking, every time a company designs an evaluation: are we measuring how this person actually works, or are we measuring how well they perform for us while being watched?

Juan Carlos Vásquez has spent ten years inside enterprise content operations, repairing content supply chains before scaling them. Fix to Flow is the discipline that work produced. The views here are his own.

Keep Reading