A single session, asked to play the role of four senior developers, just produced a working backend, a styled frontend, a security audit, and a CI/CD script in one afternoon — the kind of output that used to require a small agency and a week of meetings.
Tom’s Guide published the experiment recently, framing it as the clearest demonstration yet that one well-prompted chatbot can stand in for an entire engineering squad on scoped tasks.

Before: How Teams Actually Shipped Features
Until recently, building a single mid-size web application meant juggling calendars. A backend lead wrote the API, a frontend specialist wired up the components, a security engineer reviewed every endpoint, and a DevOps engineer handled deployment pipelines.
Tom’s Guide noted that the four-persona split mirrors how real product teams divide labor in sprint planning — each role carries distinct vocabulary, distinct priorities, and distinct blind spots. The cost of that structure is friction: handoff documents, contradictory opinions, and the inevitable Slack thread where frontend and backend argue about payload shapes.
The Catalyst: One Prompt, Four Personas
The experiment worked because accepted an explicit instruction to switch personas mid-conversation. Tom’s Guide’s writer fed the chatbot a single app brief, then asked it to respond sequentially as a backend developer, then a frontend developer, then a security reviewer, then a DevOps engineer.
Each persona came with different concerns. The backend persona pushed for clean REST contracts and idempotency keys. The frontend persona demanded accessibility tags, responsive breakpoints, and reduced-motion handling. The security persona flagged input sanitization, rate limiting, and exposed secrets in environment files.
The DevOps persona delivered a Dockerfile, a GitHub Actions workflow, and a rollback strategy. Worth noting: the personas did not always agree.
When the security reviewer insisted on stricter authentication, the backend persona adjusted the API spec on the spot — a small but telling moment that suggests can hold conflicting priorities in a single context window.
After: What a Single Chat Session Now Replaces
The after-state is that a solo developer with a good prompt can produce a first draft of an entire stack in a single sitting. Tom’s Guide’s experiment reportedly produced roughly 1,200 lines of code across four domains, plus a deployment manifest, in under an hour of interactive prompting.
That is not a replacement for a senior team on a production system, but it is a credible starting point for prototyping, hackathons, and early-stage MVPs.
Persona Output Comparison
| Persona | Primary Output | Key Concern | Sample Deliverable |
|---|---|---|---|
| Backend | API spec + database schema | Idempotency, contracts | REST endpoints with auth middleware |
| Frontend | Component tree + styles | Accessibility, responsive | React components with ARIA tags |
| Security | Threat model + fixes | Input sanitization, secrets | Rate limiter, env-file scrubber |
| DevOps | Pipeline + deployment | Rollback, reproducibility | Dockerfile + GitHub Actions YAML |
How to Use This Approach
If you are prototyping a small app and need breadth over depth, run this multi-persona prompt yourself — start with the backend brief, then explicitly request the next persona by name.
If you are shipping to production, treat the output as a scaffold for human review, not a finished system. The experiment proves can coordinate four senior voices in one session, but it does not prove it can catch the edge cases a real team would catch after a code review.
Originally reported by Tomsguide.
Related Articles
- Claude Sonnet vs GPT-4o: The Real Software Engineering Champion
- Claude Cowork Now Runs Its Own Browser Inside the Desktop App
- Claude
FAQs
What did Tom’s Guide’s experiment actually test?
It tested whether one chatbot could sequentially play four senior developer roles — backend, frontend, security, and DevOps — and produce coherent, role-appropriate output for a single app brief.
How long did the multi-persona test take?
The reported session reportedly produced roughly 1,200 lines of code across four domains in under an hour of interactive prompting.
Can really replace a four-person engineering team?
No. The experiment demonstrates useful prototyping speed, not production-grade work. Real teams still own architecture decisions, code review, and on-call responsibility.
Which persona produced the most useful output in the test?
According to the report, the security persona was reportedly the most surprising, reportedly flagging exposed environment variables and missing rate limits that the other three personas had ignored. The takeaway: ask to wear four hats on your next prototype, then hand the result to a human reviewer before shipping.
Was this article helpful?
Your feedback directly improves future articles on this site.





