How Do I Check If an Essay Was Written by ChatGPT?
The first time I seriously wondered whether an essay had been written with ChatGPT, I made the mistake of looking for a particular “AI writing style.” I expected the text to give itself away through polished sentences, predictable vocabulary, or an unusually formal tone. After reading more about how AI detection actually works, I became much less confident in that approach.
If I need to check whether an essay was written by ChatGPT, I now treat detection as an investigation rather than a simple yes-or-no test. I can use an AI detector as one source of information, but I would also compare the essay with the writer's previous work, examine the sources and citations, look at how the argument develops, and, when possible, ask the writer to explain parts of the paper. No single feature of the prose can reliably prove that ChatGPT produced it.
That matters because even OpenAI says AI detectors have not proved reliable enough to be used as the sole basis for judging students. Its guidance specifically notes that detectors can incorrectly classify human writing as AI-generated.
I stopped treating polished writing as evidence
There are certain characteristics that can make me suspicious of a piece of academic writing. Maybe the vocabulary suddenly becomes much more sophisticated. Maybe every paragraph has a remarkably similar rhythm. Perhaps the introduction makes broad claims without saying anything particularly specific about the subject.
Those observations can justify a closer look, but they don't establish where the text came from.
I have read plenty of genuinely human academic writing that sounded formulaic. Students sometimes imitate the tone of textbooks, journal articles, or model essays because they believe that is what academic writing is supposed to sound like. Someone writing in a second language may also produce unusually formal sentences because they have consciously learned academic vocabulary.
OpenAI has pointed out that AI detectors can disproportionately affect students who are learning English as a second language and can also misclassify writing that is formulaic or concise.
That changed the way I interpret stylistic clues. I might notice them, but I don't treat them as evidence by themselves.
I start with the student's writing history
If I have access to earlier assignments, this is one of the most useful comparisons I can make.
Suppose a student normally writes fairly direct paragraphs, occasionally makes grammatical mistakes, and uses a particular vocabulary. Then a new essay suddenly contains extremely polished transitions, sophisticated terminology, and a completely different sentence structure.
I would want to know why.
That difference could indicate AI assistance, but there are plenty of innocent explanations. Perhaps the student received extensive tutoring. Maybe someone proofread the paper. Maybe the assignment was edited several times. Perhaps the student simply spent much more time on this particular essay.
The comparison becomes more useful when I look beyond grammar. I pay attention to how the writer approaches evidence, explains concepts, builds arguments, and makes connections between ideas. Those habits are often more distinctive than individual words.
A sudden change across all of those areas deserves a conversation. It still doesn't prove that ChatGPT was involved.
I look closely at the sources
This is probably the part of the process I find most revealing.
AI-generated essays can contain citations that look perfectly academic while failing under closer inspection. A reference might exist but not actually support the statement attached to it. A paper may cite a real researcher who never made the claim attributed to them. In some cases, bibliographic details can also be incomplete or inaccurate.
That gives me something concrete to investigate.
I would take several claims from the essay and check them against the cited sources. If the essay says that a particular study found a specific result, I want to see whether the study actually says that. If a quotation appears in quotation marks, I want to verify the wording and location.
This is useful even when ChatGPT is not involved. Poor source checking is simply a weakness in academic writing.
I also pay attention to whether the essay contains highly specific claims supported by surprisingly vague references. A sentence might sound authoritative but lead back to a general webpage that does not contain the evidence the writer claims to have found.
For me, that is a stronger reason to investigate than a sentence merely “sounding like AI.”
I use an AI detector, but I don't stop there
There are now several tools designed to estimate whether text was likely generated by an AI system. Turnitin, for example, has an AI Writing Report that identifies qualifying text it considers likely to have originated from a large language model. Its current documentation also explicitly warns that the model may misidentify human-written, AI-generated, and AI-paraphrased text.
That limitation is important.
If I put an essay into a detector and receive a high AI-writing percentage, I treat the result as a reason to investigate further. I don't translate the percentage directly into “this student used ChatGPT.”
Turnitin itself says its AI-writing score should not be used as the sole basis for adverse action against a student. The company recommends further scrutiny and human judgment alongside institutional policies.
There is another detail I find useful when interpreting results. Turnitin currently does not surface a numerical score for results below 20 percent because it considers that range more susceptible to false-positive interpretation.
So if someone tells me that a detector gave an essay a 7 percent AI score and therefore “proved” something, I would immediately question that conclusion.
I pay attention to the argument, not just the wording
One of the better ways I have found to examine suspicious writing is to ask the writer about the actual ideas in the paper.
Imagine an essay discussing climate policy. The paper makes a sophisticated argument about the relationship between carbon pricing and household income. If I ask the writer why they chose that example, what source influenced the argument, or why one paragraph follows another, I can learn much more than I would from staring at the wording.
Someone who wrote the paper should generally be able to explain its reasoning, although I would not expect every student to give an impressive answer immediately. Nervousness, poor communication skills, or simply forgetting details after finishing an assignment can all affect a conversation.
I therefore use questions to understand the student's process rather than to conduct an improvised interrogation.
The distinction matters because academic writing is ultimately about the student's ability to understand and communicate ideas. A detector can analyze patterns in text. It cannot reconstruct the entire process by which that text came into existence.
I also check whether the essay actually answers the assignment
Sometimes the strongest clue is surprisingly ordinary: the paper doesn't quite understand the question.
AI-generated writing can produce fluent paragraphs that drift away from the precise assignment. A prompt asking for an evaluation might receive a descriptive essay. A question asking the student to compare two theories might result in two separate summaries with little comparison.
Again, that is not proof of AI use. Human students make exactly the same mistake.
But when weak alignment appears together with suspiciously generic prose, unsupported claims, questionable citations, and a major departure from the student's previous writing, I would have a much stronger reason to investigate.
I find this approach more useful because it focuses on the academic work itself. Even if no AI was involved, identifying those weaknesses tells me something valuable about the essay.
What about using an essay proofreading tool?
I would be careful here because proofreading and AI detection solve different problems.
A proofreading service can help identify grammar, spelling, punctuation, clarity, and other writing issues. It may make an essay look more polished, but that doesn't establish whether ChatGPT originally generated it.
EssayPay's educational resources and writing tools can be useful in the broader process of reviewing academic writing, particularly when I want to improve clarity or understand how a particular type of essay should be structured. I would still keep proofreading separate from questions about authorship.
This distinction is easy to lose because a polished essay can appear more suspicious simply because it contains fewer ordinary mistakes. Good editing does not automatically mean AI generation.
I would keep drafts and notes when the essay is my own
There is a practical side to this that I think students sometimes overlook.
If I write an essay myself, I would keep the outline, research notes, source files, early drafts, and revision history. These materials show how the paper developed over time. If someone later questions the authorship, being able to demonstrate the progression from notes to draft to final version is much more useful than simply saying, “I wrote it myself.”
This is especially helpful when a final version has gone through substantial editing.
A document's history cannot establish every detail of authorship, but it provides context that an isolated AI-detection score cannot. It also gives me a clearer record of my own research and reasoning.
Why I don't trust a single percentage
The biggest change in my approach has been accepting that AI detection is probabilistic.
Turnitin's current guidance describes its AI report as a tool that should support human judgment rather than replace it. Its documentation specifically says that no AI detector is perfectly accurate and that the result should be considered alongside other information.
OpenAI takes an even more cautious position, stating that its own research found AI detectors were not reliable enough for consequential judgments about students.
That doesn't mean detection tools are useless. They can identify text worth examining more closely. The mistake is treating their output as a forensic certificate.
If I were evaluating a questionable essay, I would want several independent pieces of evidence to point in the same direction before drawing a strong conclusion.
The method I would actually use
My process would start with reading the essay normally. I would first ask whether the argument makes sense, whether the evidence supports the claims, and whether the paper follows the assignment.
Then I would compare it with the student's previous writing if that material were available. After that, I would verify several citations and investigate unusual claims. If there were still concerns, I would run the essay through an appropriate AI-writing detector and interpret the result cautiously.
Finally, I would look at the writing process itself. Drafts, notes, revision history, research documents, and the student's ability to explain their own argument can all provide context.
That combination gives me a much more defensible assessment than trying to identify ChatGPT from vocabulary or sentence length.
If I had to reduce the entire process to one practical rule, I would keep it simple: an AI detector can tell me that a piece of writing deserves a closer look, but it cannot tell me the complete story of how that writing was produced. The strongest assessment comes from comparing the text with the student's established work, checking the evidence, understanding the assignment, and looking at the development of the essay over time. That approach takes a little longer, but it is far less likely to confuse polished human writing with machine-generated text.
Helpful References
How DevOps Skills Can Boost Job Opportunities for Students - Collabnix
Gamification In Education: Exploring Apps That Turn Learning Into Play
When Students Struggle Most: Peak Homework Difficulty Periods
Understanding Academic vs. Casual English in Essay Writing