Skip to main content

Still paying full price for ChatGPT Plus, Claude & Gemini?

Split the exact same subscriptions with GamsGo and cut your monthly AI bill by up to 80%, same accounts, a fraction of the cost.

See how much you save →
Sponsored
Back to the tools index

Check if my essay sounds AI generated, sentence by sentence

Paste a draft and this page measures eight writing habits, shows the number behind each one, and names the sentences carrying them. It runs in your browser and uploads nothing.

The short version

  • This measures writing style, not authorship. It cannot tell you whether a machine wrote a passage, and no honest tool can.
  • The habits it counts are the ones a marker notices anyway: identical sentence lengths, connective openers, claims attributed to nobody.
  • Detectors that promise a verdict are unreliable in a documented direction. OpenAI withdrew its own after six months.
  • Punctuation is counted and scored at zero on purpose. Swapping dashes changes nothing a reader is assessing.

Check if my essay sounds AI generated

Your draft

0 words, 0 sentences
Try one

Style score

--of 100

Not enough text to score

A score needs at least 60 words, because below that the few signals that can be measured would decide the whole result on their own.

Where the score came from

Eight measurements, weighted. Anything the passage is too short to measure is left out of the total rather than counted as zero, so 0 points of weight were in play here. The weights and the cutoffs are set by hand, not fitted to a labelled corpus.

  • Sentence length uniformitynot measured

    Needs at least 60 words across four sentences.

  • Vocabulary varietynot measured

    Needs at least 60 words.

  • Words that surged after 2022not measured

    Needs at least 60 words.

  • Stock connective openersnot measured

    Needs at least 60 words across three sentences.

  • Three part comma seriesnot measured

    Needs at least 60 words.

  • Not only, but alsonot measured

    Needs at least 60 words.

  • Sourceless attributionnot measured

    Needs at least 60 words.

  • Repeated sentence openersnot measured

    Needs at least 60 words across four sentences.

Counted but deliberately not scored

0 em dashes, 0 curly quotes, 0 en dashes and 0 unusual spaces. These are formatting artefacts, not writing habits. Swapping them changes how a draft looks and leaves every sentence saying the same thing in the same order, so they earn no points here.

Next step

If the sentences came back clean and you are still seeing odd characters in the file, the problem is formatting rather than phrasing, and that is a separate job. If the draft came out of a chat window in the first place, it is worth knowing which tools produce which habits before the next one.

What the eight signals measure

Each row is a quantity that can be counted from the text alone. Nothing on this list needs to know where the words came from, which is the whole reason the page can be honest about what it is doing. The habits were chosen for academic writing in particular, where a marker is reading for evidence and structure rather than for polish, so the attribution row matters more here than it would on a blog post.

SignalWhat is countedWhy it is on the listWeight
Sentence length uniformityStandard deviation of sentence word counts divided by the meanDrafts written in one pass tend to settle on a rhythm and hold it. Human revision leaves short sentences next to long ones.22
Vocabulary varietyUnique words per rolling 50 word windowA rolling window is used because a plain unique word ratio always falls as a text gets longer, which would make short drafts look better than long ones for no reason.16
Words that surged after 2022Hits from a fixed word list, per thousand wordsWords such as delve, underscore and pivotal became more common in academic abstracts after chat models arrived, measured across a corpus rather than per paper. One hit means nothing. Full weight is awarded at twelve per thousand words, a cutoff picked by hand.18
Stock connective openersShare of sentences starting with Moreover, Furthermore, Additionally and their relativesThese are the joins a model reaches for when it has no argument connecting two paragraphs, only an order.14
Three part comma seriesComma series of exactly three parts, per thousand wordsThree is the list length that arrives when there is no fourth item in mind. A four item list contains a three item tail, so a match sitting behind a comma is discarded. The pattern cannot tell noun phrases from coordinated clauses, so both are counted.8
Not only, but alsoNegative parallelism constructions, per thousand wordsThe construction adds emphasis without adding a claim, so it accumulates in text optimised to sound significant.8
Sourceless attributionStudies show, experts agree and similar, in a sentence carrying no citation markerA sentence that names its source, by year in parentheses, by a bracketed reference number or by et al, is left alone. This is the one a marker checks first, and it costs marks whether or not a machine was involved.8
Repeated sentence openersShare of sentences whose first word opens another sentence tooA cheap proxy for paragraph structure that keeps resetting to the same starting position.6

The word list behind the third row overlaps the excess vocabulary reported in a Science Advances analysis of biomedical abstracts (preprint here), which tracked which words became more frequent after chat models arrived by comparing against pre 2022 baselines across millions of abstracts. That analysis works at the level of a corpus and explicitly cannot say which individual abstract was affected, so treat the word list as a description of a trend rather than as evidence about your draft.

Why detectors get this wrong, and in which direction

The failure is not random. Detectors that report a probability are mostly reporting how predictable your writing is to a language model, and predictability tracks fluency. Liang and colleagues put 91 TOEFL essays by non native English speakers through seven detectors and more than 61 percent came back labelled machine generated, against a correct human call on more than 90 percent of essays by eighth grade students in the United States. The summary of that work is blunt about the mechanism: a narrower band of common words produces lower perplexity, and lower perplexity is what gets read as machine.

The company with the best access to the models withdrew from this. OpenAI launched an AI text classifier in January 2023 that correctly identified 26 percent of machine written text while calling human text machine written 9 percent of the time, and shut it down that July citing its low rate of accuracy. Six months, from launch to withdrawal, by the organisation best placed to make it work.

Institutions have done the arithmetic since. Vanderbilt disabled the Turnitin AI detector for its instructors in August 2023, and the reasoning is worth reading in full: at the one percent false positive rate the vendor claimed, the roughly 75,000 papers Vanderbilt submitted to Turnitin in 2022 would have produced around 750 papers incorrectly flagged as machine written. A rate that sounds small stops sounding small once it is multiplied by a cohort.

That is the context this page sits in. The query people type, check if my essay sounds AI generated, usually lands them on a tool that returns a percentage, and the percentage is the part the evidence above does not support. What you get here instead is numbers you can check yourself, which is a smaller promise and a keepable one.

What to change in a draft that scores high

Start with the sourceless attribution rows, because those cost marks on their own terms. Every sentence the panel flagged for studies show or experts agree is a claim you have asserted without naming who made it. Find the source or drop the claim. This is the only category where the fix is not stylistic.

Then delete the connective openers rather than replacing them. A paragraph beginning with Moreover almost never needs the word, and cutting it forces you to check whether the paragraph actually follows from the one above it. Sometimes it does not, which is the useful discovery.

For sentence length, the lever is a single deliberate short sentence per paragraph. Not a rewrite. The measurement is a spread around a mean, so one genuinely short sentence placed where the argument turns moves it further than nudging a run of sentences into slightly different lengths, and it reads better besides, because that is what emphasis is for.

Leave the punctuation alone. The panel counts your em dashes and awards them nothing, and the reason is in the arithmetic above: nothing that assesses your work is looking at them. Time spent swapping dashes is time not spent on the two paragraphs that needed a source.

How the scoring works

Sentence splitting runs first, on a whitespace normalised copy, breaking at a terminator followed by a space and a capital or an opening quote. Abbreviations that end in a period are held back by a list, and a single letter before a period is treated as an initial, so a citation like J. Zou does not become two sentences. Word counts come from a Latin letter tokeniser, which means numerals and symbols do not inflate a sentence length.

Each of the eight signals produces a raw measurement and a value between zero and one describing how far that measurement sits toward the templated end. That value is multiplied by the signal weight to give points. The total is points earned over points available, not points earned over 100, and the difference matters: a passage too short to measure vocabulary variety has that weight removed from the denominator rather than scored as zero. Otherwise every short paste would score artificially well.

The per sentence rows use the same functions as the aggregate, and a row is dropped when its signal was not measurable, so a row can never claim something the total was not allowed to count. Two of the eight signals, length uniformity and vocabulary variety, are properties of the whole passage and have no per sentence equivalent, which is why the rows list six flag types rather than eight.

The weights and the cutoffs are chosen by hand. Nothing here was fitted against a labelled corpus of machine written and human written essays, because no such corpus exists that I would trust, and a scorer calibrated on a bad one would be worse than an honest arbitrary one. Read the individual measurements, which are facts about your text, ahead of the composite, which is my weighting of them. Below sixty words no score is reported at all, since one or two measurable signals would otherwise decide the whole result.

Everything runs against a deferred copy of your text, so typing into a long draft stays responsive, and the sentence rows stop rendering past 400 sentences while the counts continue to cover the whole thing. There is no network request in any of it, which you can verify yourself rather than take on faith.

What this will not tell you

It will not answer the question in the title. Asked as a question about style, check if my essay sounds AI generated has a useful answer and this page gives it. Asked as a question about authorship it does not, because sounding like something and being it are separate claims and only the first is measurable from the text. It measures habits that are common in machine output and also common in rushed human writing, formal institutional prose, and the work of anyone taught to signpost every paragraph. Those populations overlap heavily, and no amount of counting separates them.

It will not predict a detector score. Those tools work on token level predictability, which this page does not compute and could not compute without sending your draft to a model. If a score here and a score there disagree, that is expected, not a bug in either.

It has a bias worth naming. Writing that uses short, common words scores worse on vocabulary variety, which means it will tend to be harder on second language writers in the same direction the commercial detectors are, for a related reason. The third gallery card exists to show that in the open rather than bury it. Read the score as a description of the prose, and give the vocabulary row less weight than the attribution row if English is not your first language.

The attribution check is sentence level and imperfect in both directions. A sentence carrying any citation marker is exempt in full, so a cited claim sitting beside an uncited one is missed, and a source named without a year or a bracket, as in a bare surname, is not recognised as a source at all. It errs toward saying nothing, which is the right direction for a tool nobody should be accused on the strength of.

The word list is fixed and will age. The vocabulary that marked machine output in 2024 is already drifting as models change and as writers adapt away from it. A word list compiled from published abstracts describes the past, and this one gets no automatic updates.

Frequently asked questions

How can I check if my essay sounds AI generated without uploading it?

Paste it into the box above. The measurement runs inside the page, so the draft never leaves your machine and there is no account, no queue and no copy of your work sitting on a server. You can confirm that in about twenty seconds by opening your browser network tab, pasting a paragraph, and watching that no request goes out. That check is worth running on any tool that asks for an unpublished draft, this one included.

Does a high score here mean a detector will flag my essay?

No, and the two questions are less related than they look. This page counts writing habits you can see and edit. A commercial detector reports how predictable your word choices are to a language model, which is a different quantity with a different failure mode. A draft can score high here because it leans on connectives and still pass a detector, and a draft can score low here and still be flagged. Treat the score as an editing checklist rather than a prediction.

Why do AI detectors flag essays by non native English speakers?

Because fluency and predictability are hard to tell apart with these methods. Liang and colleagues, writing in Patterns and linked in the detectors section above, ran 91 TOEFL essays by non native speakers through seven detectors and more than 61 percent came back labelled as machine generated, while essays by eighth grade students in the United States were correctly called human more than 90 percent of the time. Their explanation is that writing drawn from a narrower band of common words is easier for a model to predict, and predictability is what the detectors are ultimately measuring. Under that account the tools are scoring vocabulary range, not authorship.

Does my essay sound like ChatGPT?

That depends on which habits it picked up, and the panel above names them one at a time rather than answering yes or no. The pattern people recognise as ChatGPT in academic writing is usually four things at once: sentences of near identical length, a connective at the top of every paragraph, three item lists where the argument wanted two or five, and claims handed to studies with no study named. Any one of those is ordinary. All four together is the texture people are reacting to when they say a piece reads like a chatbot wrote it.

Is an em dash a sign of AI writing?

It is the tell people repeat most and it is the weakest one on the list. Chat output does use em dashes more heavily than most human drafts, but that is a formatting habit rather than a watermark, and plenty of careful writers use them constantly. This page counts them and deliberately awards them zero points, because replacing every dash leaves your sentences saying the same thing in the same order, which is the part anyone is actually assessing.

What does burstiness mean in AI writing detection?

Burstiness is the informal name for variation in sentence length and structure across a passage, as opposed to perplexity, which describes how surprising the next word is to a language model. This page measures the first one directly, as the standard deviation of sentence word counts divided by the mean, and reports the number rather than hiding it inside a verdict. It does not measure perplexity, because doing that honestly requires running a language model over your text, which would mean uploading it.

My essay was flagged as AI. What can I actually do about it?

Evidence of process is the thing that has worked for students, and it is the thing you can prepare now: version history in the document, dated notes, a draft with the crossings out still in it, and the sources you read. Institutions have been backing away from detector output as sole evidence for this reason. Vanderbilt disabled the Turnitin AI detector for its instructors in August 2023, pointing out that a claimed one percent false positive rate would still mean around 750 incorrectly flagged papers across the roughly 75,000 it submitted to Turnitin in 2022. Its statement is linked in the detectors section above.

Can I make AI writing sound human by swapping words?

Not by swapping words alone, no. Synonym substitution leaves sentence lengths, paragraph shapes and the order of your claims exactly where they were, and those are the properties this page measures. What changes the reading is structural work: cutting a sentence in half, deleting the connective and letting two sentences sit next to each other, replacing a generic claim with the specific thing you read. That is also the work that improves the essay, which is the more useful reason to do it.

About the author

Jim Liu

Jim Liu runs the OpenAI Tools Hub review portfolio. He built this after reading enough detector marketing to notice that none of it would state which quantity was being measured, and deciding that a page which shows its arithmetic was more useful than a fifteenth tool that shows a percentage. More about how this site tests things.

Sponsored

Ad served by Adsterra. OpenAIToolsHub is not responsible for advertiser content.