If you want to work on a serious Agent / Harness team, the strongest portfolio signal is the ability to explain why an agent behaved a certain way and run an experiment that changes that behavior.
Two external signals
Vlad Feinberg's frontier lab hiring advice separates two adjacent opportunities: lower-level kernel work under the LLM stack, and agentic loops above the LLM abstraction. The latter asks for rigorous, controlled, technical experiments around LLM agent behavior.
DeepSeek Harness hiring adds the product context. Tianyi Cui's formula is Model + Harness = Agent. The model sets the capability ceiling; the harness decides how that capability enters
context, tools, tasks, evals, user feedback, and training signal.
Together, the message is practical: top Agent / Harness teams need evidence that you can turn agent behavior into something observable, comparable, and reproducible.
Demos are weak evidence
Most agent demos have three gaps:
- Too few tasks, so the demo proves one run.
- No baseline, so attribution becomes guesswork.
- No trace, so failure analysis becomes guesswork.
A small controlled experiment carries more signal. Take the same task set and compare a baseline prompt, a tool whitelist, typed errors, a reviewer subagent, and a compaction policy. Then show success rate, cost, latency, failure categories, and a few traces. The project can be plain; the point is control over harness variables.
A portfolio should read like an experiment report
A portfolio project for Agent / Harness roles should answer six questions:
| Question | Artifact |
|---|---|
| What should the agent do? | task set / scenario definition |
| What is the baseline? | baseline result |
| What changed? | intervention |
| How is quality measured? | eval metrics / graders |
| Where did failure happen? | trace / diagnostics |
| Can someone rerun it? | scripts / fixtures / README |
This matches Anthropic's agent eval guidance: agent evaluations often grade both transcript and outcome, using a mix of code-based, model-based, and human graders. Final answers alone miss tool choice, retries, state pollution, and benchmark loopholes.
Pick narrow experiments
Good portfolio work studies one variable well inside a real agent runtime.
Context experiment
Compare full transcript, rolling summary, retrieved memory, and branch summary on long-task success. This maps to Memory Systems and Claude Code's Context.
Tool experiment
Compare coarse tools with fine-grained tools, free-form errors with typed errors, long observations with compressed observations. Anthropic's tool guidance is clear: tool response structure can change eval performance. This maps to Tools & Actions, MCP & Skills, and OpenClaw's Tools.
Trace / eval experiment
Define a trace schema: message, tool_call, tool_result, state_change, diagnostic, final_outcome. Then combine code graders, LLM judges, and human checks. This maps to Evaluation and Claude Code's Debugging.
Multi-agent experiment
Compare single agent, planner-executor, planner-executor-reviewer, and parallel subagents. The point is to explain when decomposition improves success and when it adds cost or coordination failure. This maps to Multi-Agent Systems, Claude Code's Subagents, and OpenClaw's Multi-Agent Routing.
Production / safety experiment
Study permission policy, hooks, human checkpoints, rollback, and cost caps. This maps to Production, Agent Security, and OpenClaw's Security.
A concrete project template
Project:
agentic-debug-harness
Question:
When a coding agent fixes failing tests, when does it over-edit unrelated files?
Task set:
20 small bugfix tasks. Each task includes the initial repo, failing tests, expected behavior, and allowed edit scope.
Baselines:
- naive prompt
- prompt + file whitelist
- prompt + permission hook
- prompt + reviewer subagent
- prompt + trace replay
Metrics:
- test pass rate
- unrelated file edits
- token / cost
- wall-clock latency
- human intervention count
- failure taxonomy
Deliverables:
tasks.jsonlrun-eval.tstraces/*.jsonlresults.csv- five short failure analyses
- reproducible README
The project matters because it shows how a harness design changes behavior.
The most valuable skill is attribution
Strong candidates can separate failure sources:
- The model missed the task.
- The tool schema encouraged misuse.
- The observation was too long or too vague.
- Context lost key state.
- Memory stored a false fact.
- Subagent handoff dropped a constraint.
- Permission policy was too loose or too strict.
- The eval only measured surface success.
That judgment is difficult to prove in a resume. A trace, a small intervention, and a before/after result prove it much better.
A 90-day route
Weeks 1-2: reproduce
Reproduce one agent task: coding, tool-use, or research. Get the task, runner, trace, and result table working.
Weeks 3-5: change one variable
Pick context, tool schema, hook, memory, or subagent design. Avoid changing prompt, model, and tools at once; the experiment should remain interpretable.
Weeks 6-8: add traces and evals
Record each run as an episode. At minimum, capture tool calls, tool results, stop reason, and final outcome. Track success, cost, and failure type.
Weeks 9-12: write the report
Write it like an experiment report: question, baseline, intervention, task set, metrics, trace examples, conclusion.
Bottom line
Agent usage is table stakes. Agent / Harness teams need people who can put agent behavior on a test bench: control variables, record process, compare outcomes, and explain failures.
One rigorous small experiment is stronger evidence than ten polished demos.