A real-world coding-agent challenge · Research report 01
Kill My SaaS. Keep the workflow.
Agents can build a convincing interface. Replacing the software someone depends on means understanding the entire job.
We examined 69 reviewed submission records, the wider builders’ process accounts, and three finalist interviews with the people who run the work. The most useful lessons live between the specification, the generated product, and the operator trying to use it.
Evidence cutoff: October 11, 2026. This is a process and evaluation report. A winner, final numerical scores, and prize allocation are not established in the evidence reviewed.
Open-call speaker → review
Invited speaker → onboarding
Sponsor → assigned deliverables
→
A role-specific portal
Complete deliverables
A usable, live schedule
The operational target, schematically. These paths do not all begin with an application form.
01 / The evidence
A small field study. A demanding real job.
The challenge was to build useful conference operations software: collect submissions, review them, onboard speakers, collect deliverables, communicate, and manage a schedule.
The brief called for conditional or routed CFP forms, a self-service speaker portal, templated communications and calendar integrations, evaluation and scoring, drag-and-drop scheduling with conflict handling, and an onboarding dashboard. An open-source repository and a deployed, testable site were requested. Usefulness to the AI Engineer team mattered; exact visual cloning was not the objective. [B]
72raw form records
69mapped to the review roster
3finalist interviews
66other reviewed records
Records are the unit, not unique people. Duplicate identities, rebrands and resubmissions exist. Matching submission IDs excludes three raw records from this roster; this is reconciliation, not an inferred disqualification. [R]
Builders combined tools
Among the other 66 records, 48 selected more than one coding tool and 59 selected more than one model family. Forty selected both Claude Code and Codex. Tool selections describe reported use, not time, token volume, or effectiveness. [S]
Other 66 records
48 / 66
Selected multiple coding tools.
Other 66 records
59 / 66
Selected multiple model families.
These counts do not rank the models. All three finalists selected the form’s “GPT 5.6 Sol” label; so did 50 of the other 66. Builders chose their own tools, combined them, and reported them imperfectly. Form labels are preserved as recorded, including aliases; they are not independently verified commercial model identities. [S]
The finalists described different ways to organize the work. Their accounts share a practical starting point: construct a product model from evidence before asking agents to implement it.
Process descriptions below combine builder self-reports and October 10 interviews. Named reviewer observations are positive highlights, not comprehensive pass verdicts. These accounts do not establish which method caused finalist selection.
Marko Kraemer
Trackstage · Interview M
Video & documentation
Product and API map
Durable delegation
Adversarial review
Marko’s written account describes an orchestrator, two to nine subagents, persistent task and agent documents across context compaction, and review through additional coding tools. In the interview he describes taking the supplied recording, documentation and API reference, choosing a familiar stack, and having the agent map the API. [M 14:34–15:18; S]
“I let the agent self-map out the entire API reference.”
The written review praises clean drag-and-drop and clean UX. The transferable technique is to make the domain and integration surface explicit enough that delegated work can refer back to a stable source. Agent count by itself is not a quality measure. [R]
Timing caveat: the written account says roughly 24 hours; the call says roughly 48. The scopes are unresolved. An anecdote about rebuilding an earlier construction project is a separate project, not a measured Trackstage speedup.
Valentin De Matos
OpenRostrum · Interview V
Explore screens & flows
Markdown specifications
Dependency-aware workers
Improve the factory
Valentin framed his experiment as discovering how far the process could be automated. He describes using Codex to reconstruct flows from web material and video frames, writing feature expectations in Markdown, then deciding which tasks could run concurrently. His written process adds static lint rules, worker lanes and specialist critics. [V 03:27–05:24; S]
A short manual user journey remained part of the method. Around 07:07–08:28, he describes investigating traces and missing context, then improving the research and onboarding inputs so the process could be rerun. That is a useful account of process repair, rather than a claim of zero human involvement.
Kelsey’s interview feedback praises the multiple-portal work; the written review also calls out neat UI and clean drag-and-drop. The lesson to carry forward is that role and workflow coverage can be a visible product strength. [V 05:24–06:32; R]
Yazin
Conferencer · Interview Y
Research help material
Clarify the specification
Click through personally
Refine the interface
Yazin describes converting the brief to Markdown, using Firecrawl to gather documentation and related product material, and spending a couple of hours clarifying the specification. His written account describes planning followed by implementation in Cursor and iterative feedback. [Y 01:45–03:34; S]
He then personally navigated the product and maintained a fix list. He describes reorganizing menus, icons, and table-heavy pages to make the underlying tasks easier to understand. This is product editing as a distinct phase of agent-assisted development. [Y 03:41–06:40]
Operator payoff · Y 05:36–06:05
Kelsey praised seeing deliverable statuses together on one page, and relayed a colleague’s positive reaction. The improvement reduced the need to open individual forms to understand readiness. This is a paraphrase of reviewer feedback, not a directly recorded quotation from that colleague.
The written 4.5-hour run and oral five-to-six-hour estimate are approximate and do not include all research and human iteration. Neither is a verified total time to replace the incumbent.
Useful methods beyond the finalists
Bodo reports converting 42 screenshots into screen-level specifications and encoding recurring agent mistakes as lint rules or edit-time hooks. Sessionbored describes research spikes, a 29-table domain model, worktrees and simulated-persona tests. Opensesh describes giving workers separate branches, databases, ports and environment files, with one integration owner and manual walkthroughs of the deployed product. These are positive process examples from submission narratives, not independent endorsements of overall product quality. [S]
03 / Interactive workflow snapshot
Follow the person, not just the screen.
A form can collect information. An operational product also needs to know who the person is, how they arrived, what they owe, and what changes when their state changes.
The interviews distinguish open-call, invited-speaker and sponsor paths. Explore the implications below. This is an illustrative reconstruction for evaluation design, not a screenshot or live test of any entrant. The scenarios and checks are our recommendations derived from the reviewed workflows.
Conference operations / Scenario explorerSynthetic example · no participant data
Completion needs observable consequences
An acceptance button should have consequences for portal access, messaging, deliverables and schedule readiness. A role switch should preserve the correct permissions and make the way back understandable. The evaluator needs to inspect those consequences, not merely confirm that a control exists.
This is our synthesis of the brief, reviewer notes and interviews. It is not a newly measured universal cause of failure. [B, R, M, V, Y]
04 / The review perspective
Where operational coverage became visible.
Manual coding of notes across the 69 reviewed records found repeated mentions of portal limitations, scheduling, narrow form-only coverage and access problems. The categories overlap. These are counts of mentions in review notes, not exhaustive feature tests or precise failure rates.[R]
Access observations comprise seven dead links and ten authentication, expired-access or admin-access issues. Inaccessibility does not prove a feature never existed. A later shutdown cannot be retroactively attributed to the judging date.
Move a session, handle a conflict, see the resulting state
Seed a conflict and record the resolution
A demo with privileged access
Clear organizer and participant access
Fresh-session tests for each role
A polished empty state
A short path to the first useful event
Observe a new organizer’s first five minutes
A passing isolated feature
Consistent effects across the workflow
Check data, notifications and views after a change
The right-hand columns are recommended evaluation practice, not claims that every listed test was performed during this challenge.
Review categories are not final scores
The recorded labels are 3 Finalist, 2 Strong, 12 Maybe, 29 Weak, 22 No and 1 blank. They sum to 69 and describe the sheet’s categories. They do not establish a final ranking, numerical weighting, winner or prize decision. Individual negative notes remain outside this public report.
05 / For builders of large agent projects
Make the project understandable and testable.
The useful unit of delegation is a piece of the product with inputs, boundaries, and a way to tell whether it works.
Research the domain before dividing the code. Combine help articles, video walkthroughs, API references and operator questions. Write roles, states, transitions and exceptions into the specification. A screenshot describes one state, not the rules that produced it.
Use durable working documents. Keep requirements, decisions, task status and unresolved questions in files that survive context resets. Require workers to identify changed assumptions and integration risks in their handoffs.
Parallelize along explicit dependencies. Establish shared identity, data contracts and permissions first. Isolate branches and development data where appropriate. Give one owner responsibility for integration.
Evaluate the running environment. Launch the app, use fresh sessions, test role boundaries and verify the deployed system. An agent’s completion report or demonstration video is a useful artifact, not a replacement for observed task completion.
Fix the process when a class of errors repeats. Turn recurring mistakes into lint rules, acceptance checks or improved research inputs. Preserve the failing scenario as evidence.
Budget human product work explicitly. Distinguish research, agent execution, review, debugging and operational maintenance. Comparing only a fast implementation run obscures the work that made it useful.
Recommendations synthesized from submission narratives and the finalist interviews. These practices recur in the evidence; their individual causal effects were not isolated. [S, M, V, Y]
“It's about thinking through the prompt.”
In that exchange Kelsey asks whether the builder understands why a feature is requested and the journey around it. Yazin responds in terms of owning the software and considering downstream effects. That is a requirement for product responsibility, not merely better prompt syntax. [Y 15:19–16:44]
06 / For organizers of similar projects
Design the evaluation as carefully as the brief.
A useful challenge should leave behind reproducible evidence about the work, even when it cannot produce a controlled model comparison.
Before submissions
Version the governing brief and distinguish core requirements, optional features and bonuses.
Provide realistic personas and fixtures: open-call, invited and sponsor paths; an incomplete deliverable; a scheduling conflict.
Publish the access contract: test credentials, seeded data, expiry dates, deployment retention and a route for reporting evaluator access failures.
Ask for process artifacts as well as code: specs, prompts, decision logs, agent handoffs and a declared human-work budget.
During review
Freeze the evaluated revision, deployment URL and observation time. Keep later improvements in a separate record.
Capture task outcomes and short evidence clips. Distinguish passed, failed, inaccessible and not tested.
Combine repeatable checks with an operator walkthrough. Include the first five minutes and a state-changing task.
Give feedback tied to the attempted task, with confidence and scope. Retain entrant-level criticism privately where appropriate.
What the community wanted to learn
In the bounded Discord material reviewed, participants wanted reusable workflows, prompts and session histories, plus actionable evaluation feedback. There was interest in sharing learnings that could support future projects and client work. These are qualitative themes, not a poll or a complete server census. [D]
A practical response is to publish the evaluation protocol and aggregate findings alongside builder process accounts. Keep reimbursement documentation and results communication distinct. Historical claims about rule exceptions require the governing brief or a documented amendment; community recollection alone does not settle eligibility.
Perceived absence of evaluation traffic also does not establish that an entry was skipped. That question requires deployment mapping and evaluation logs. A contest build and a subsequently improved product should never silently become the same observation.
This challenge suggests research directions around specification recovery, long-horizon coordination, and independent evaluation. These are hypotheses and proposed experiments, not findings of a controlled benchmark.
Can an agent recover the missing domain model?
A brief leaves many roles, exceptions and transitions implicit. Test whether an agent can identify the missing information, ask useful questions and represent it coherently.
Proposed experiment: hold the implementation budget constant; compare brief-only inputs with documentation, operator Q&A and state diagrams. Measure completion on unseen role journeys.
Does a local fix preserve the whole workflow?
Changes to acceptance, identity or scheduling affect several surfaces. A patch can satisfy the immediate request while creating inconsistent downstream state.
Proposed experiment: issue a realistic mid-project change and check invariants across permissions, notifications, deliverables and schedule output.
How do we evaluate an evaluator?
An agent can optimize toward its own interpretation of a rubric. An internal perfect score is not equivalent to independent user success.
Proposed experiment: separate builder and evaluator contexts; hold back scenarios; compare claimed completion with fresh-session operator observations. Record disagreement and access failures explicitly.
What survives a long run?
Parallel work and context resets put pressure on shared requirements and integration decisions. Durable documents help only if agents keep them accurate and consult them.
Proposed experiment: introduce handoffs and context resets at controlled points; measure contract drift, integration rework and unacknowledged assumption changes.
Can usefulness be measured beyond feature presence?
A consolidated deliverables view can save effort without adding a new business capability. Navigation and information layout affect whether people can use a feature at all.
Proposed experiment: measure task time, recovery from mistakes and operator intervention on equivalent capabilities with different interaction designs.
Who owns the system after the demo?
The interviews bring maintenance and accountability into the definition of success. Software replacement includes handling change and failure after launch.
Proposed experiment: extend evaluation into a maintenance period with incident diagnosis, a policy change and data repair; track human intervention as well as code output.
08 / Selected interview moments
A guide to the process conversations.
Filter these public-safe selections by conversation. Summaries are editorial paraphrases; timestamps are elapsed time in the original October 10 recordings. They are approximate source windows for later editing, not frame-accurate cuts. Private recordings are not embedded here.
Marko: local offline ASR without diarization; named attribution is contextual. Valentin and Yazin: original Zoom speaker-labeled transcripts. None of the excerpts on this page has been checked by direct listening. Machine errors and speaker-label errors remain possible.
09 / Sources, scope & limits
What this evidence can support.
This report consolidates a submission analysis, review-note coding, bounded community reading and three interviews. The publication contains curated aggregates and excerpts; raw submissions, contacts, private source links, full recordings and individual negative reviews are not included.
[B] Challenge brief
Core functional requirements and delivery expectations, as reconciled in the source analysis. Struck optional items included AI review across rounds, Accelevents integration, a resource/wiki portal and embedded views. No numerical weighting was recovered. Governing historical versions and exceptions remain unresolved for definitive eligibility claims.
[R] Review roster and notes
69 mapped records in the judging sheet, matched by submission ID to the raw responses. Review-category totals are reported without inventing scores. The four thematic counts are manually coded, overlapping mentions in written notes. They are not a new product audit.
[S] Builder submissions
72 raw records; 69 in the reviewed cohort and 66 outside the finalist group. Tool and model selections are self-reports. Narratives may mention tools omitted from selection fields; charts preserve the selections instead of silently augmenting them.
[M] Marko / Trackstage
October 10, 2026, 3:30 pm EDT interview. Recovered recording: 27:22.72; 499 offline ASR segments. Original elapsed timing was preserved. No diarization; attribution requires care.
[V] Valentin / OpenRostrum
October 10, 2026, 4 pm EDT interview. Original Zoom transcript: 163 cues through 19:38.580. Labels and text may contain transcription errors.
[Y] Yazin / Conferencer
October 10, 2026, 9 pm EDT interview. Original Zoom transcript: 158 cues through 23:24.319. Reviewer remarks and builder accounts are distinguished from editorial inference.
[D] Community discussion
Bounded visible Discord reading: announcements spanning August 8–October 10 and general discussion in late September/early October. Coverage was not exhaustive. Themes are qualitative, not prevalence estimates; displayed message timezones were not confirmed.
Why there is no cost, speed or model leaderboard
Time accounts mix active work, research, wall-clock duration and agent runtime. Spend fields mix subscriptions, API equivalents, currencies and other units. The participants selected their own tools and used multiple models. Interviews are retrospective self-reports. A pooled average or causal ranking would imply comparability that the evidence does not provide.
What would strengthen the next study?
Frozen build hashes and deployment snapshots; complete timestamped task logs; versioned rubrics; independent operator trials; recorded human intervention; comparable budgets; maintained access; and clearly separated contest, interview and later product versions. Audio verification is still required before treating the selected transcript excerpts as verified spoken quotations or final subtitles.
Prepared with AI assistance from the source inventory and transcript analysis. Editorial synthesis and proposed evaluation recipes are labeled separately from observations. This page reports the available evidence as of October 11, 2026 and can be updated when final decisions or stronger verification are available.