← Insights
Perspectives · Jun 23, 2026 · 7 min read

Preparing for Agent/Harness teams: portfolio to experiments

EssayCareerEngineering

If you want to work on a serious Agent / Harness team, the strongest portfolio signal is the ability to explain why an agent behaved a certain way and run an experiment that changes that behavior.

Two external signals

Vlad Feinberg's frontier lab hiring advice separates two adjacent opportunities: lower-level kernel work under the LLM stack, and agentic loops above the LLM abstraction. The latter asks for rigorous, controlled, technical experiments around LLM agent behavior.

DeepSeek Harness hiring adds the product context. Tianyi Cui's formula is Model + Harness = Agent. The model sets the capability ceiling; the harness decides how that capability enters context, tools, tasks, evals, user feedback, and training signal.

Together, the message is practical: top Agent / Harness teams need evidence that you can turn agent behavior into something observable, comparable, and reproducible.

Demos are weak evidence

Most agent demos have three gaps:

  1. Too few tasks, so the demo proves one run.
  2. No baseline, so attribution becomes guesswork.
  3. No trace, so failure analysis becomes guesswork.

A small controlled experiment carries more signal. Take the same task set and compare a baseline prompt, a tool whitelist, typed errors, a reviewer subagent, and a compaction policy. Then show success rate, cost, latency, failure categories, and a few traces. The project can be plain; the point is control over harness variables.

A portfolio should read like an experiment report

A portfolio project for Agent / Harness roles should answer six questions:

QuestionArtifact
What should the agent do?task set / scenario definition
What is the baseline?baseline result
What changed?intervention
How is quality measured?eval metrics / graders
Where did failure happen?trace / diagnostics
Can someone rerun it?scripts / fixtures / README

This matches Anthropic's agent eval guidance: agent evaluations often grade both transcript and outcome, using a mix of code-based, model-based, and human graders. Final answers alone miss tool choice, retries, state pollution, and benchmark loopholes.

Pick narrow experiments

Good portfolio work studies one variable well inside a real agent runtime.

Context experiment

Compare full transcript, rolling summary, retrieved memory, and branch summary on long-task success. This maps to Memory Systems and Claude Code's Context.

Tool experiment

Compare coarse tools with fine-grained tools, free-form errors with typed errors, long observations with compressed observations. Anthropic's tool guidance is clear: tool response structure can change eval performance. This maps to Tools & Actions, MCP & Skills, and OpenClaw's Tools.

Trace / eval experiment

Define a trace schema: message, tool_call, tool_result, state_change, diagnostic, final_outcome. Then combine code graders, LLM judges, and human checks. This maps to Evaluation and Claude Code's Debugging.

Multi-agent experiment

Compare single agent, planner-executor, planner-executor-reviewer, and parallel subagents. The point is to explain when decomposition improves success and when it adds cost or coordination failure. This maps to Multi-Agent Systems, Claude Code's Subagents, and OpenClaw's Multi-Agent Routing.

Production / safety experiment

Study permission policy, hooks, human checkpoints, rollback, and cost caps. This maps to Production, Agent Security, and OpenClaw's Security.

A concrete project template

Project:

agentic-debug-harness

Question:

When a coding agent fixes failing tests, when does it over-edit unrelated files?

Task set:

20 small bugfix tasks. Each task includes the initial repo, failing tests, expected behavior, and allowed edit scope.

Baselines:

  • naive prompt
  • prompt + file whitelist
  • prompt + permission hook
  • prompt + reviewer subagent
  • prompt + trace replay

Metrics:

  • test pass rate
  • unrelated file edits
  • token / cost
  • wall-clock latency
  • human intervention count
  • failure taxonomy

Deliverables:

  • tasks.jsonl
  • run-eval.ts
  • traces/*.jsonl
  • results.csv
  • five short failure analyses
  • reproducible README

The project matters because it shows how a harness design changes behavior.

The most valuable skill is attribution

Strong candidates can separate failure sources:

  • The model missed the task.
  • The tool schema encouraged misuse.
  • The observation was too long or too vague.
  • Context lost key state.
  • Memory stored a false fact.
  • Subagent handoff dropped a constraint.
  • Permission policy was too loose or too strict.
  • The eval only measured surface success.

That judgment is difficult to prove in a resume. A trace, a small intervention, and a before/after result prove it much better.

A 90-day route

Weeks 1-2: reproduce

Reproduce one agent task: coding, tool-use, or research. Get the task, runner, trace, and result table working.

Weeks 3-5: change one variable

Pick context, tool schema, hook, memory, or subagent design. Avoid changing prompt, model, and tools at once; the experiment should remain interpretable.

Weeks 6-8: add traces and evals

Record each run as an episode. At minimum, capture tool calls, tool results, stop reason, and final outcome. Track success, cost, and failure type.

Weeks 9-12: write the report

Write it like an experiment report: question, baseline, intervention, task set, metrics, trace examples, conclusion.

Bottom line

Agent usage is table stakes. Agent / Harness teams need people who can put agent behavior on a test bench: control variables, record process, compare outcomes, and explain failures.

One rigorous small experiment is stronger evidence than ten polished demos.

Related paths