How to Tell If Code Is AI-Generated (and What to Ask)
You cannot tell reliably. In published tests, AI detectors applied to source code performed poorly, and the signs reviewers rely on, such as commit metadata and style patterns, can be removed or imitated. The more useful question is whether the code was reviewed and tested, because review and testing are what find the documented problems in AI-written code.
The security record of AI-generated code is covered in a companion article: Vibe coding security: what the research and the incidents show.
What is AI code generation?
AI code generation is the use of a large language model to produce source code from a natural-language request or from the surrounding code. It takes three common forms: autocomplete inside an editor, chat that returns code on request, and agents that plan a task, edit files and run commands.
Most developers use these tools or plan to. In the 2025 Stack Overflow Developer Survey, "84% of respondents are using or planning to use AI tools in their development process". Sundar Pichai said on October 29, 2024 that "more than a quarter of all new code at Google is generated by AI, then reviewed and accepted by engineers."
For a codebase written in the last two years, the likely answer to "is this code AI-generated" is therefore "partly".
Can you tell if code is AI-generated?
Not reliably. Peer-reviewed studies and the vendors of detection tools say the same thing.
| Source | Date | What was tested | Finding |
|---|---|---|---|
| Pan et al. | January 8, 2024 (ICSE-SEET 2024) | Five detectors, including GPTZero, Sapling and DetectGPT, on 5,069 Python problems and ChatGPT's solutions | "existing AIGC Detectors perform poorly in distinguishing between human-written code and AI-generated code" |
| Suh et al. | November 6, 2024 (ICSE 2025) | GPTZero, Sapling, GPT-2 Output Detector, DetectGPT, GLTR and the code-specific GPTSniffer, on code from four models | "they all perform poorly and lack sufficient generalizability to be practically deployed" |
| OpenAI | January 31, 2023 | Its own AI text classifier | "it is unreliable on code". Withdrawn "As of July 20, 2023" because of "its low rate of accuracy" |
| Turnitin | Read October 3, 2026 | Its AI writing detection | "does not reliably detect AI-generated text in the form of non-prose, such as poetry, scripts, or code" |
The detail supports the summaries. Suh et al. write of the general-purpose detectors that "their Accuracy is mostly less than 0.6", where 0.5 is a coin toss. The classifier the authors trained themselves reached "an F1 score of 82.55", which is better and still short of proof for any single file.
Research detectors keep improving. Shi et al., first posted January 12, 2024 and presented at ICSE 2025, proposed a detector, DetectCodeGPT, and reported that it beat earlier methods in their experiments. The same paper opens by noting that language models "blur the distinctions between machine- and human-authored source code".
Three limits apply to all of this research:
- The studies used models from 2023 and 2024, and mostly short functions from benchmark sets.
- A detector trained on one model's output must generalize to models released later.
- Code that a person edits after generation is a mix, and no study above claims to separate the two parts.
A wrong result has a cost. OpenAI reported that its text classifier labeled human-written text as AI-written "9% of the time". Applied to a vendor's or an employee's code, an error rate like that means false accusations.
What signs do reviewers actually look for?
Reviewers look at metadata first, then at style, then at mistakes that models are known to make. Each sign is circumstantial.
| Sign | Where to look | Why it is weak evidence |
|---|---|---|
| Tool trailers and bot authors in commits | git log, pull request history |
Trailers can be switched off or stripped |
| Agent instruction files | Repository root and tool folders | Shows a tool was set up, not which lines it wrote |
| Concise, uniform code with a narrow vocabulary | The code itself | Formatters and style guides produce the same effect |
| Explicit error handling everywhere | The code itself | Careful engineers write this too |
| Naming or formatting that shifts between files | Across modules | Common in any codebase with several authors |
| A dependency that does not exist | Package manifests | Rare from people, but a typo can cause it |
Metadata. Coding tools mark their work by default. Claude Code's documentation, read on October 3, 2026, states that "Commits get a git trailer such as Co-Authored-By by default", and documents a setting to "Change or hide the trailer Claude Code adds to commits". GitHub's documentation, read the same day, says: "Commits from Copilot cloud agent are authored by Copilot, with the person who started the task listed as co-author."
Researchers depend on these markers and state the limit. Georgia Tech reported on April 13, 2026 that its Vibe Security Radar "can trace metadata like co-author tags, bot emails, and other known tool signatures, but it can't identify an issue if these markers have been removed." Researcher Hanqing Zhao added: "Claude Code and Copilot together account for most of what we detect, but that's partly because they leave the clearest signatures."
Instruction files. Many agents read a project file of instructions. The AGENTS.md site, read on October 3, 2026, describes "a simple, open format for guiding coding agents" that is "used by over 60k open-source projects".
Style. Shi et al. compared human and machine code and found that "Machines tend to write more concise code", and that machines prefer "a limited spectrum of frequently-used tokens, whereas human code exhibits a richer diversity in token selection." They also found that "Machine-authored code pays more attention to exception handling and object-oriented principles than human". These are averages across the datasets and models the authors tested. They do not settle the origin of one file.
Review findings. CodeRabbit reported on December 17, 2025, from 470 open-source pull requests, that "Readability issues spiked more than 3× in AI contributions" and that "AI introduced nearly 2× more naming inconsistencies". CodeRabbit sells an AI review tool, and it could not verify authorship: "Since it was impossible to directly confirm authorship of each PR of a large enough OSS dataset, we checked for signals that a PR was co-authored by AI".
Invented dependencies. Spracklen et al., first posted June 12, 2024 and presented at USENIX Security 2025, generated 576,000 code samples and found "the average percentage of hallucinated packages is at least 5.2% for commercial models and 21.7% for open-source models". An import of a package that has never existed is the closest thing to a fingerprint, and it is also a security risk.
Why do people ask whether code is AI-generated?
People ask because of three concerns: security, maintainability and ownership. Each has published evidence behind it.
Security. Veracode reported on July 28, 2026 that "The average security pass rate across models is 56%" on its security tasks. Results varied by flaw type: models passed SQL injection tasks 83% of the time and cryptography tasks 87%, but "performance fell sharply on cross-site scripting at 15% and log injection at 12%." Veracode sells security testing, and the figures are pass rates on tasks built to test security.
Maintainability. He et al., first posted November 6, 2025 and accepted at MSR '26, compared open-source projects that adopted Cursor with matched controls. They found "a statistically significant, large, but transient increase in project-level development velocity, along with a substantial and persistent increase in static analysis warnings and code complexity."
Ownership. The US Copyright Office's report on copyrightability, published January 29, 2025, concludes that "Copyright does not extend to purely AI-generated material, or material where there is insufficient human control over the expressive elements." It also concludes that "The use of AI tools to assist rather than stand in for human creativity does not affect the availability of copyright protection for the output." How this applies to a given codebase is a question for counsel, and it matters in acquisitions and licensing.
A fourth reason is misplaced confidence. In a user study by Perry et al., first posted November 7, 2022 and published at ACM CCS 2023, participants with an AI assistant "wrote significantly less secure code than those without access" and "were more likely to believe they wrote secure code". The model in that study is outdated, and the confidence effect is the finding that lasts.
What is the more useful question to ask?
The more useful question is whether the code was reviewed and tested, and by whom. The evidence supports this for three reasons.
Authorship cannot be established. The detectors are unreliable and the markers are optional, as shown above.
The risks are properties of the code. A missing access check or unescaped output fails the same way whoever typed it, and both can be found by testing the code directly.
Teams that use AI at scale pair it with review. Google's sentence ends with "then reviewed and accepted by engineers." He et al. conclude: "Our study identifies quality assurance as a major bottleneck for early Cursor adopters". Google's 2025 DORA report, announced September 24, 2025, puts it as "AI doesn't fix a team; it amplifies what's already there."
Developers apply the same caution. In the 2025 Stack Overflow survey, "More developers actively distrust the accuracy of AI tools (46%) than trust it (33%)".
For work delivered by a vendor, disclosure works better than detection. Ask in the contract which AI tools are used, require that tool trailers stay on, and require human review of every change.
How do you review code regardless of who wrote it?
Check nine things. Each maps to a failure documented in the research above, and none depends on knowing the author.
| # | Check | What to look for |
|---|---|---|
| 1 | Explanation | A named engineer can explain each module and why it is built that way |
| 2 | Review history | Every change went through a pull request that a second person approved |
| 3 | Tests | Automated tests exist, run on every change, and cover the paths that touch money and personal data |
| 4 | Access control | Authorization is enforced on the server; test it logged out and as a different user |
| 5 | Input handling | Forms and endpoints withstand hostile input, including script tags and log injection |
| 6 | Secrets | No keys in client code or in the git history |
| 7 | Dependencies | Every package exists, is the one intended, and has no known vulnerability |
| 8 | Complexity | Static analysis runs in CI, and warnings are tracked over time |
| 9 | Error paths | The code handles a failed call, a timeout and an empty result |
Checks 4 and 5 target the classes where Veracode found models weakest. Check 7 follows Spracklen et al. Check 8 follows He et al. Check 9 follows CodeRabbit's finding that "Error handling and exception-path gaps were nearly 2× more common" in the AI-co-authored pull requests it reviewed.
Key takeaways
- No detector identifies AI-generated code dependably. Two ICSE studies found existing tools "perform poorly", and OpenAI and Turnitin say their text detectors are unreliable on code.
- Commit trailers and agent files show that a tool was used. They can be removed, and their absence proves nothing.
- Style patterns exist on average and do not settle the origin of a single file.
- The documented risks, such as weak input handling, rising complexity and invented dependencies, can be tested directly.
- Ask who reviewed the code and which tests it passes. For vendors, require disclosure of AI tool use in the contract.
Frequently asked questions
Is this code AI-generated?
For a single file, no tool gives a dependable answer. Check the commit history for tool trailers such as Co-Authored-By and the repository for agent instruction files. If they are absent, authorship stays unknown, and the practical step is to review and test the code.
Are AI code detectors accurate?
Not on current evidence. Suh et al. (ICSE 2025) tested six detectors and concluded "they all perform poorly and lack sufficient generalizability to be practically deployed". Their own trained classifier reached an F1 score of 82.55 on their benchmark data.
Can GPTZero or Turnitin detect AI-generated code?
Turnitin states that its model "does not reliably detect AI-generated text in the form of non-prose, such as poetry, scripts, or code". GPTZero was among the detectors that Pan et al. and Suh et al. tested on code, and both studies reported poor results for the tools they tested.
Is AI-generated code worse than human-written code?
It depends on the task and on review. Veracode's July 2026 report put the average security pass rate at 56%, with 83% on SQL injection tasks and 15% on cross-site scripting. CodeRabbit counted 10.83 issues per AI-co-authored pull request against 6.45 for human-only ones, using inferred authorship.
Does it matter legally whether code is AI-generated?
It can. The US Copyright Office concluded on January 29, 2025 that copyright "does not extend to purely AI-generated material", while assistive use of AI "does not affect the availability of copyright protection for the output." Ask counsel how that applies to your code, especially before a sale or a licensing deal.
Easital Technologies Ltd. reviews codebases on these terms, whoever or whatever wrote them. See code audit services for an independent review and vibe coding rescue for repairing an AI-built application before it goes to production.
Sources
All sources were opened and checked on October 3, 2026.
- Stack Overflow, 2025 Developer Survey, AI section, 2025. https://survey.stackoverflow.co/2025/ai
- Sundar Pichai, "Q3 earnings call: CEO's remarks", Google, October 29, 2024. https://blog.google/inside-google/message-ceo/alphabet-earnings-q3-2024/
- Pan, Chok, Wong, Shin, Poon, Yang, Chong, Lo, Lim, "Assessing AI Detectors in Identifying AI-Generated Code: Implications for Education", arXiv 2401.03676, January 8, 2024 (ICSE-SEET 2024). https://arxiv.org/abs/2401.03676
- Suh, Tafreshipour, Li, Bhattiprolu, Ahmed, "An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far Are We?", arXiv 2411.04299, November 6, 2024 (ICSE 2025). https://arxiv.org/abs/2411.04299
- OpenAI, "New AI classifier for indicating AI-written text", January 31, 2023, with a notice of withdrawal dated July 20, 2023 (read through a reader proxy). https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
- Turnitin, "Using the AI Writing Report", guide, undated (read through a reader proxy). https://guides.turnitin.com/hc/en-us/articles/22774058814093-Using-the-AI-Writing-Report
- Shi, Zhang, Wan, Gu, "Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human Programmers", arXiv 2401.06461, January 12, 2024 (ICSE 2025). https://arxiv.org/abs/2401.06461
- Anthropic, Claude Code documentation, "All settings" (attribution settings), undated. https://code.claude.com/docs/en/settings-reference
- GitHub, Copilot documentation, "Tracking GitHub Copilot's sessions", undated. https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent/track-copilot-sessions
- Georgia Tech, "Bad Vibes: AI-Generated Code is Vulnerable, Researchers Warn", April 13, 2026. https://news.research.gatech.edu/2026/04/13/bad-vibes-ai-generated-code-vulnerable-researchers-warn
- AGENTS.md project site, undated. https://agents.md/
- CodeRabbit (David Loker), "Our new report: AI code creates 1.7x more problems" (State of AI vs Human Code Generation Report), December 17, 2025. https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report
- Spracklen, Wijewickrama, Sakib, Maiti, Viswanath, Jadliwala, "We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs", arXiv 2406.10279, June 12, 2024 (USENIX Security 2025). https://arxiv.org/abs/2406.10279
- Veracode, "2026 GenAI Code Security Report: AI Is Writing More of Your Code but Security Hasn't Caught Up", July 28, 2026. https://www.veracode.com/blog/2026-genai-code-security-report-ai-risk/
- He, Miller, Agarwal, Kästner, Vasilescu, "Speed at the Cost of Quality: How Cursor AI Increases Short-Term Velocity and Long-Term Complexity in Open-Source Projects", arXiv 2511.04427, November 6, 2025 (MSR '26). https://arxiv.org/abs/2511.04427
- US Copyright Office, "Copyright and Artificial Intelligence, Part 2: Copyrightability", January 29, 2025. https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf
- Perry, Srivastava, Kumar, Boneh, "Do Users Write More Insecure Code with AI Assistants?", arXiv 2211.03622, November 7, 2022 (ACM CCS 2023). https://arxiv.org/abs/2211.03622
- Google Cloud, "Announcing the 2025 DORA Report", September 24, 2025. https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report

