OpenAI launched GPT-5.3-Codex on February 5, 2026, marking a decisive shift from code assistant to comprehensive digital coworker. The model doesn’t just write code—it handles product documentation, data analysis, spreadsheets, and multi-day projects while running 25% faster than its predecessor, positioning Codex for the full spectrum of professional computer work.
GPT-5.3-Codex Launch: Key Features & Capabilities
Release Overview
| Feature | GPT-5.3-Codex | GPT-5.2-Codex |
|---|---|---|
| Speed | 25% faster | Baseline |
| SWE-Bench Pro | 56.8% | 56.4% |
| Terminal-Bench 2.0 | 77.3% | 64.0% |
| OSWorld-Verified | 64.7% | 38.2% |
| Cybersecurity Rating | High (first model) | — |
| Availability | Paid ChatGPT plans | API |
| Launch Date | February 5, 2026 | January 2026 |

Beyond Code: Full Professional Workflows
GPT-5.3-Codex represents OpenAI’s vision of an end-to-end digital collaborator. The model participates across the entire software lifecycle—debugging, deploying, monitoring, writing product requirement documents, editing copy, designing tests and metrics, creating presentations, and managing spreadsheets. This expansion transforms Codex from a coding companion into a general-purpose professional assistant.
OpenAI emphasizes the model can handle long-running tasks lasting hours or even multiple days while maintaining context and accepting mid-task steering without losing prior decisions. Users can interact in real-time, asking questions and guiding the solution as the model provides frequent progress updates.
Benchmark Leadership & Efficiency
The model achieves record performance on key benchmarks while using fewer output tokens than predecessors, potentially lowering costs per task. The strongest improvements appear on terminal-driven and computer-use tasks (OSWorld-Verified jumping from 38.2% to 64.7%), indicating significant gains in hands-on system work beyond synthetic tests.
While the SWE-Bench Pro improvement is incremental (56.8% vs 56.4%), OpenAI reports fixes for real engineering pain points including lint loops, weak bug explanations, and premature completion on flaky tests.
Self-Improving AI: Building Itself
In a remarkable development, GPT-5.3-Codex helped build itself—early versions debugged the training pipeline, managed deployment, and diagnosed test results. This self-hosted capability accelerated development but raises questions about transparency and audit trails when AI systems shape their own toolchains.

First “High” Cybersecurity Rating
GPT-5.3-Codex is OpenAI’s first model classified as “High capability” for cybersecurity under its Preparedness Framework. While OpenAI lacks “definitive evidence” the model can fully automate cyberattacks, they’re deploying their most comprehensive safety stack including automated monitoring, trusted access programs, and enforcement pipelines.
CEO Sam Altman acknowledged the model could “meaningfully enable real-world cyber harm, especially if automated or used at scale.” API access is delayed while OpenAI implements safeguards, with security features including the Aardvark security research agent and free codebase scanning for open-source projects like Next.js.
Competitive Timing
The launch occurred just 72 hours after Anthropic’s Claude Opus 4.6 release, intensifying the agentic AI race. Both companies are racing to expand beyond narrow coding tasks toward comprehensive professional automation, though OpenAI’s cybersecurity concerns highlight the risks of increasingly powerful autonomous systems.
For more AI development news, visit TechnoSports AI.





