Writing text with artificial intelligence is simple. Determining whether text is written by a person or by artificial intelligence is quite difficult.
An AI Detector is a computer program that examines text and estimates the probability that the text was generated by a given artificial intelligence system, for example, ChatGPT, Claude, Gemini, or other artificial intelligence systems.
AI Detectors can be useful to screen various types of text including essays, articles, job applications, product reviews, and others. It is important to keep in mind that an AI Detector can only provide an estimate. There is no way for the Detector to prove beyond a reasonable doubt who is the actual author of the text. There is always a possibility that an AI Detector makes an error.
Turnitin, which is a plagiarism detection system, states that its AI writing detection system should not be used to take adverse action against a student.
In general, an AI Detector should be viewed as an indication that text is artificial intelligence generated text and not as a definite proof.
Some of the things tools may look for when analyzing text include:
- How words and phrases are used
- How sentences are structured
- How consistently and predictably language is used
- How different forms of structure and style vary
- How text is repeated
- What statistical distributions exist in text
- How text is structured at the sentence level
- Text that contains patterns that are known to occur in text generated by AI
- Edits and changes to text made by AI
Often, different approaches are used. Sapling, for example, describes a system that performs token- and sentence-level analyses. GPTZero distinguishes its work from other systems by stating that it uses an end-to-end deep learning approach anddoes not rely on older metrics like perplexity and burstiness.
Generally, text is analyzed in order to identify patterns. Patterns, along with other factors, are then used to estimate the probability that AI was involved in text’s creation.
What Does an AI Detector Actually Detect?
AI detectors can’t identify an “AI signature” that is invisible to humans.
AI detectors look for signs that may indicate that text was created by a computer.
Examples of signs that may be looked for may include:
- Use of extremely generic and vague transitional phrases
- Use of overly simple and/or overly complex sentence structures
- Use of sentence structures and/or sentence lengths that are consistently the same throughout a piece of text
- Use of prose that is overly consistent
- Use of writing structures that are similar to writing styles or patterns used to train the AI detector
These signs may also be present in text that is written by a human.
AI writing may be edited or paraphrased to remove AI-like writing structures, styles, and/or prose.
Therefore, classification of text as being AI-written may not be 100% accurate.
How Do AI Detectors Work?
Although different AI Detector systems may differ in their specific methods, most systems may use similar processes to classify text.
1. The detector receives the text
Users copy or upload text or supported document types. Some services can extract AI-written text from documents. GPTZero and Turnitin, for example, can process different document types.
2. The system analyzes linguistic aspects
Various elements of writing can be analyzed.
Some examples are:
- Frequencies of words and phrases
- Word and sentence order
- Style
- Type of vocabulary
- Variety
- Repetition
Previously, explaining how detectors work focused more on AI aspects, such as burstiness and perplexity.
Burstiness is the variation in the structure of different sentences and phrases in a text.
Perplexity, on the other hand, is the level of surprise or how unpredictable a text is to a language model. If a text is highly unpredictable to a model, it means that the text does not contain a lot of useful information for the model to learn.
However, burstiness and perplexity should not be relied on for detection, and in some cases, they may lead to a false positive.
3. The model makes a classification
Detection models are developed to identify various aspects in writing, and are trained to distinguish between human-written and AI-written texts.
Sapling and GPTZero are some examples. Sapling states it employs deep-learning architectures to perform sentence-level classification and token-level scoring to assess probability, while GPTZero states it uses deep-learning models to perform the same tasks.
4. The Detector Issues a Classification or Probability
A hypothetical output may read:
Probability of Human Input: 0.7
Probability of AI Input: 0.3
Mixed
AI-generated text
AI-enhanced text
Probability of Human Input: 0.15
Probability of AI Input: 0.85
The terminology used to represent output varies by tool.
For example, one tool may represent a 90% score as “this output is likely generated by ChatGPT,” while another tool may represent the same score as “extremely high confidence in ChatGPT output.”
A score of 90% does not always represent a strong confidence level in AI output. This is often overlooked.
Can AI Detectors Detect ChatGPT?
Several AI detectors have been built to identify outputs generated using ChatGPT, Google Bard/Gemini, Microsoft Bing AI (Copilot), DeepSeeking, Scribm, LLaMA, and similar models.
Scribbr, for example, states its AI Detector can identify outputs generated by Gemini and Copilot, in addition to ChatGPT. Similarly, Copyleaks asserts its AI Detector can identify outputs generated by Gemini, Claude, and other models in addition to ChatGPT.
It is also important to note that the ability of an AI detector to reliably identify a particular version of an AI model does not guarantee that the detector will be able to reliably identify all future models.
Similar to AI systems, the systems designed to detect AI systems are also continually improving and changing.
How Reliable Are AI Detectors?
100% accuracy is not achievable for any AI detector.
This is the most crucial element to understand when utilizing an AI checker.
Different implementations of a tool may result in varying output for the same input. Additional factors which may impact performance include:
- Length of the input
- Input language
- Style of the input
- Topic of the input
- Selected AI model
- Level of human post-processing
- Degree of paraphrasing
- Translations
- Level of human and AI collaboration in content creation
- A detector’s prior experience with a sample of generated text
Weaknesss in an AI’s abilities to model different styles of human writing have been documented. The 2025 NAACL paper assessing text deceitation states that, “…in some cases, a high false positive rate may be required, and even then some samples may still evade detection.”
The 2024 NIST GenAI pilot study evaluates AI text detection by framing the issue as one of distinguishing AI-generated text from human-generated text. The study suggests that in many cases, perfect differentiation of human from machine generated text cannot be expected.
There are many ways to unintentionally introduce stylistic elements that may lead a detector to conclude text was AI generated. For example, academic and formal writing, and writing that is repetitive and/or overly concise may lead to a detector concluding that text was generated by an AI.
Turnitin recognizes that AI text detection may result in false positives. It has attempted to address this by adjusting its threshold for interpreting AI text detection. Scores below Turnitin’s threshold are not reported to the user.
It can especially happen when:
- There is substantial editing
- There is paraphrasing
- It is translated
- It is created using different models from those used by the AI detector
- It is of short or strange structure
- There are mixed human inputs
AI deters usually fail when there is humanization and paraphrasing.
Can AI Detector Models Determine Authorship?
They cannot.
Two different things are being asked here.
An AI Detector can assess the probability that a given piece of text was created by AI.
To determine authorship one must observe how the author creates the text e.g. the author’s editorial process, drafts, and writing history.
This is why, in addition to text, evidence of the writing process (e.g. drafts, notes, editing history) and other pertaining to the case (e.g. citations, institutional policies, etc.) should be assessed when reviewing text for potential academic integrity violations.
Turnitin says that for academic integrity cases, human judgment should be relied upon and and its AI should be viewed as a tool to assist.
AI detectors and plagiarism checkers fulfill distinct functions.
Like plagiarism, AI content detection software is still in development and is not infallible. Scribbr’s AI Detector and Turnitin’s AI-Writing Percentage are separate from their plagiarism detection tools.
So why do some classified human-written contents raise suspicions about AI?
Formulaic Writing
If a document has sentence structures repeated throughout, AI-type structures may be presumed.
Written like a phone simulator?
Let us not forget the strict 20 character limit.
Polished Writing
Editing may remove irregularities and unusual structures in a sentence, similar to AI-written structures.
Academic Writing
May be predictable and strict, similar to AI.
Short Writing
Less information equals fewer structures for a classifier to evaluate, similar to AI.
Non-native English Writing
Per recent research, some AI detectors raise concerns about false positives in graduated student writing.
Because AI detectors remain flawed and improving AI content generation software is trending, detecing AI generated content remains a challenge.
Just because a user edits AI generated content, does not mean the content will no longer be classified as AI generated.
The most essential thing to know about the following types of content is:
- Human-written content
- Content that was written by AI and then edited by people.
Also, there is:
- Content that was written by people and then edited by AI.
Differentiating these types of content is something that automated systems are developed to do. For example, Scribbr talks about content that is completely generated by AI, content that is edited by people with the help of AI, and content that is written entirely by people.
The type of assistance that is provided by AI to a content author can range from minor editing to complete content authoring.
How to Use an AI Detector
When reviewing content, think of an AI detector as just one of the tools you would use.
To get a meaningful result, a detector needs to be presented with a sufficiently long input. For example, the suggested minimum length of input that one detector requires is 250 words.
Review the content that the detector highlighted. If the detector listed examples, review those.
To verify content authorship, you should be able to identify the differences between previously verified content that is written by that same person.
Contact a person who is able to review changes in the content and give you an edtiorial history of the document.
Thoroughly review the content and identify any potentially false or misleading statements. Also check the references to see if they support the content or if they are outdated.
Sapling and Turnitin both note that their AI detectors should not be the sole basis to take a negative personnel action.
How to Evaluate AI Detectors
There are numerous AI detection solutions, and you should assess each of them based on the following factors.
1. Coverage of AI models
Does the solution provide details of its plan for covering newly released AI models?
2. False positives
Have the vendors published information regarding false positives?
3. Information disclosure
Do they share information regarding the factors that impact the detection score?
4. Supported languages
If your content is in different languages, do the available solutions cover them?
5. Length of content
Do solutions perform well with short content?
6. Trust
Do they share information regarding what happens to the content once it’s submitted?
7. Automated system replacement
Do the solutions support human judgment?
A few of the most widely-used AI detectors are listed below.
- Turnitin AI Writing Detection
- Pangram
- Winston AI
- ZeroGPT
- Sapling AI Detector
- Copyleaks AI Detector
- QuillBot AI Detector
- Scribbr AI Detector
- Originality.ai
- GPTZero
These vendors have not disclosed their underlying AI models and databases. Thus, they should not be compared to one another regarding accuracy. When evaluating these vendors’ solutions, you should review their whitepapers regarding solution testing to understand their accuracy claims.
What happened to OpenAI’s AI Classifier?
OpenAI made a text classifier that purportedly identified text written by artificial intelligence (AI).
OpenAI removed this classifier on July 20, 2023, from public access due to low accuracy. OpenAI disclosed that this classifier achieved an accuracy of only 26% in identifying text written by different AI systems, and also falsely identified 9% of text written by humans as being generated by AI.
This situation is historically significant as it demonstrates the limitations of utilizing an AI detector to answer the question “Was text written by AI?,” and provides an example of why such question should not be posed to an AI detector.
Evolving Nature of AI Detection
Like other technology, the field of AI detection will also continue to evolve and change.
Detection of AI-generated text will also continue to improve.
Additionally, researchers will also continue to identify and publish the blind spots that occur when AI text detectors come across modified or advanced AI text.
This means the accuracy of a given AI detector will not be indicative of its accuracy in the future.
How to Understand Scores Obtained from AI Detectors?
The focus should not be on determining with certainty whether text was written by AI.
Rather, the text in question should be analyzed to determine if it contains any patterns of AI text, and what additional evidence should be evaluated to make a determination.
Reasons to Use an AI Score
- Evaluate large amounts of content quickly
- Pinpoint content requiring review
- Enable quality review of content
- Provide insight into use of AI in content review process
- Explore unusual/unexplained content
- Use an appropriate explanation/rationale if thestakes are high.
AI Detector FAQ
What is an AI Detector?
An AI Detector is a piece of software that estimates the likelihood of a given piece of text being produced or modified by AI.
Do AI Detectors always work?
No. AI detectors are susceptible to generating false results. Numerous variables impact the performance of AI detectors and a given text’s estimation, including the language, editing, and the AI and detection frameworks used.
Can AI detectors identify text generated by ChatGPT?
Many AI detectors are trained to identify text generated by ChatGPT and other large language models. However, there are no guarantees regarding other AI frameworks and text generations.
Can an AI detector be used to prove plagiarism?
No. AI detectors and plagiarism frameworks serve different purposes.
Can a human-written text be misidentified as AI-written text?
Yes. False positives are possible.
Can text generated by AI be misidentified as human text?
Yes. Text generated by AI can be misclassified, especially if it has been edited or modified.
How long should my input be?
The input should be as long as necessary to give the tool the information it needs to generate its output. In general, long inputs are preferable to short inputs. However, some tools have restrictions on the length of the inputs they can process. In these cases, you should use the shortest input that captures the essentials. The tools Scribbr and Sapling suggest that longer inputs can help the tools provide a more accurate output.
Are AI detectors the same as plagiarism checkers?
No. Plagiarism checkers and AI detectors assess different things. AI detectors can help identify passages in a document that may have been written by AI, but plagiarism checkers can not help identify passages written by AI.
What is your take on AI detectors?
Writing is a blend of the author’s style and the voice of the AI. In the case of technology blending the human and artificial intelligence, it becomes difficult to determine where one ends and the other begins. AI detectors help identify blurred lines. AI detectors highlight segments of writing that may raise concerns for a human.
AI detectors will always be limited to providing an assessment of a document and will not be able to conclusively prove that a certain person authored the document. In doubtful situations, AI detectors should be used in conjunction with traditional methods, such as assessing the document’s history and the author’s intent, and looking for other evidence that may corroborate the document’s authorship.
We must manage our expectations when it comes to evaluating AI. It is better to gain a deeper understanding of why we got a particular score, and evaluate the context in which that score was assigned.











