Enter a page URL or paste your own text. The counter separates the article from the menus, footer and banners, counts both with OpenAI's own tokenizers in your browser, and splits the article at its headings so you can see how big every section is.
One fetch, two texts. For a URL, our server fetches the page once as VerandBot/1.0, 12 second timeout, without running JavaScript. It returns the article, meaning the first <article> element, else <main>, else the body minus header, nav, footer and aside, with its headings kept, plus every visible word of the page. Each text is capped at 100,000 characters.
In your browser, with gpt-tokenizer 2.9.0, a JavaScript port of OpenAI's published encodings, pinned and integrity-checked. The count is exact for the text under that encoding. A chat request adds a few formatting tokens of its own, which aren't included. Sections split at H1 to H3; Short (under 40 tokens) and Long (over 800) are Verand's rule of thumb, not any engine's.
No Claude or Gemini number. Anthropic and Google don't publish tokenizers you can run locally for their current models, so we don't print one or guess a ratio. Images, video and markup aren't counted, and there's no price, because token prices change with every release.
One fetch, two texts, one tokenizer, then the article cut at its headings. The card is the tool in motion on an example guide, looped, and each step lights up while the card is doing it.
Our server fetches the URL a single time and reads the HTML without running JavaScript. Pasted text skips this step entirely: it never leaves your browser.
The same rule Verand's site audit uses on every page finds the article region and drops menus, header, footer and sidebars. Everything visible is kept as a second text, so the two can be compared.
OpenAI's o200k_base encoding runs on your machine and cuts the text into the pieces a GPT model reads. “Depreciation” alone is three of them. Switch to cl100k_base for GPT-4 and GPT-3.5.
Each H1 to H3 starts a section, and each section gets its own count. One over 800 tokens is marked Long, one under 40 is marked Short: a rule of thumb for passages that stand on their own.
Most token counters answer one question: how many tokens is this text? A page owner has two more. How much of my page is the article and how much is the template around it, and does each section hold up when an assistant lifts it out on its own?
A token is the unit a language model reads and writes. Before a model sees any text, a tokenizer cuts it into pieces drawn from a fixed vocabulary: whole common words, fragments of rarer words, punctuation, digits, and the spaces in front of words. The model never sees letters or words as such; it sees a sequence of token numbers. Context windows, rate limits and API bills are all measured in these pieces, which is why a word count is the wrong unit for any of them.
Tokens are not words, and the gap is widest exactly where regulated firms write. Run through OpenAI's current tokenizer, “Depreciation recapture is the part of a gain taxed when you sell.” is 16 tokens for 12 words: depreciation alone becomes Dep reci ation, and recapture becomes rec apture. Numbers split hardest: 1031 is two tokens, $150,000 is four, and a year such as 2025 comes out as 202 and 5. Names, acronyms, code and non-English text cost more tokens per word than plain English prose does.
So there's no honest fixed ratio of words to tokens. Two of the pages checked with this tool on 29 September 2026 show the spread: Willowdale Equity's home page article came out at 2,060 words and 2,733 tokens, 1.33 tokens per word, while Verand's own home page, dense with product labels and figures, ran at 1.52. The card prints the ratio for the text you give it, measured, rather than a rule of thumb.
OpenAI publishes its tokenizers, which is what makes an exact count possible outside its servers. Its own library is called tiktoken; this page uses gpt-tokenizer, a JavaScript port of the same encodings that runs entirely in your browser, pinned to one version and loaded with an integrity hash so the file can't change underneath the page. Two encodings matter today:
| Encoding | Used by | On this page |
|---|---|---|
o200k_base | GPT-4o, GPT-4.1, the o-series and the GPT-5 family | The default, because it covers OpenAI's current models |
cl100k_base | GPT-4 and GPT-3.5 Turbo | A toggle on the card, loaded only if you use it |
The names describe the vocabulary: roughly 200,000 pieces against roughly 100,000. A larger vocabulary holds more whole words, so the same text usually comes out slightly shorter on o200k_base. On Willowdale's home page the article was 2,733 tokens on o200k_base and 2,775 on cl100k_base. For English prose the two stay close; for other languages the difference can be larger, which is one reason the card lets you switch.
The count is exact for the text under that encoding. It is not quite what a chat request is billed, because the API wraps each message in a few formatting tokens of its own, and a request also carries its instructions and the model's reply. For sizing a page, the text count is the number that matters.
Anthropic and Google don't publish tokenizers you can run locally for their current models. Anthropic offers a free counting endpoint in its API instead, and its documentation calls the result an estimate that can differ by a small amount from what a message actually uses. The same documentation says Claude Opus 4.7 and later models use a newer tokenizer that produces approximately 30 percent more tokens for the same input than earlier models, and tells developers to recount rather than reuse old figures. Google's API has a token-counting call of its own.
That rules out the shortcut some counters take, which is to run an OpenAI tokenizer and print the result under a Claude label, sometimes with an accuracy range attached. No published source supports a fixed conversion, and a single ratio would be wrong for one Claude generation or the other. So this card shows the counts it can compute exactly and says plainly which ones it can't.
Paste-in counters count whatever you paste. A web page is harder, because what a visitor reads as “the article” sits inside a template: navigation, a header, cookie banners, related-post rails, a newsletter form, a footer with legal disclosures and a link to every section of the site. All of it is text, and all of it is tokens.
The card separates the two with the rule Verand's site audit uses on every page it crawls. It takes the page's first <article> element if it holds real text, otherwise <main>, otherwise the body with its header, nav, footer and aside removed, and drops scripts, forms and embedded frames. That's the article. Separately, it keeps every visible word of the served page. Both are counted with the same tokenizer, and the bar on the card shows what share of the page's tokens is the article.
The two pages we checked sit at opposite ends. Willowdale's home page is 91 percent article: 2,733 of its 2,997 tokens. Verand's own home page is 22 percent: 3,657 of 16,686. The second number is also a lesson in reading the result. Verand's home page carries animated replicas of the product, and the first <article> element on the page is one of them, so the “article” counted there is a product view rather than the page's copy. When the share looks wrong, read the opening lines the card shows; they tell you which region was picked, and whether your template's markup points at the text you think it does.
Why it matters: an assistant that fetches your page and passes it to a model hands over text, and a template-heavy page spends tokens on menus and disclosures before it reaches the argument. How much of that a given assistant keeps isn't published, so no count predicts whether a page gets used. What the card gives you is the proportion, which is the part you control.
A context window is the most text a model can take in at once, counted in tokens, and it has to hold the instructions, the question, any retrieved pages and the answer together. Window sizes change with almost every model release, so the card doesn't name models. It shows your article and your whole page against generic sizes, 8K, 32K, 128K, 200K and 1M tokens, and you check the model you care about against its own documentation.
A single article rarely troubles a modern window: Willowdale's home page is under 3,000 tokens. Where the question bites is combination. An assistant answering from several sources, a support bot grounded in your whole knowledge base, or you pasting five competitor pages into one prompt all add up fast, and a page that is four-fifths template fills the budget with the wrong text. The smaller windows still matter too, for the fast, cheap models that often do the retrieval and summarising in a larger system.
Systems that answer from documents rarely use a whole page. They split it into chunks, index those, and retrieve the few that match a question; that's the idea behind retrieval-augmented generation. Chunking strategies vary: fixed token lengths with some overlap, splits at structural boundaries such as headings and paragraphs, or splits where the meaning shifts. None of the major assistants publishes how it chunks the pages it reads.
So the chunk view on this card is labelled for what it is: an illustrative rule, not any engine's. It cuts the article wherever an H1, H2 or H3 starts, counts each section, and flags two cases. A section under 40 tokens, a heading with a line under it, has little to say if it's retrieved alone; link lists and calls to action often look like this and are fine where they are. A section over 800 tokens, roughly 600 words, probably covers more than one question, and any system that cuts it into smaller pieces will separate claims from the context that explains them. Willowdale's home page shows the second case: its “Investor questions, answered” section is 1,287 tokens, a whole FAQ under one heading, where each question under its own subheading would stand as its own passage.
The reason headings matter is quotability. When an assistant lifts a passage, the passage has to make sense without the paragraphs around it. A section that opens under a heading naming its question, answers it directly, and stays on that question is the easiest thing to quote accurately. The AI Content Readiness Checker grades a page on that directly; this card shows you the section sizes behind it.
Most AI token counters, including the GPT tokenizer and ChatGPT tokenizer tools that fill search results, are built for prompts: paste text, get a number, sometimes a price. That covers the paste box here too, with the token visualizer on the card showing where each piece begins and ends. The URL mode is for the questions a site owner asks instead: how heavy is my template compared with my copy, how long is this guide in the unit a model reads, and which sections would a retrieval system have to cut. The Website Word Counter answers the same questions in words; tokens are the unit the model reads.
Six things that are true of this tool, each one backed by a line in the code that runs it.
A URL count sends one address and nothing else. There's no account, no session and no database behind the tool, so there's nothing for us to keep about you.
Paste mode makes no request to Verand at all. The tokenizer runs on your machine, so a draft under embargo or a client document can be counted without leaving it.
OpenAI publishes its encodings, and the card runs them, pinned to one library version and checked against an integrity hash. The same text gives the same count every time.
The article region is found by the same extraction Verand's site audit runs on every customer page, called directly. The text you see counted is the text the product works from.
No Claude or Gemini figure, no invented ratio, no price. The chunk rule is labelled as ours, and the 100,000-character text limit is stated on the card when a page reaches it.
A URL count is one page fetch with paid renders switched off, and the counting happens on your machine, so it costs nothing and is never metered. The one limit is a courtesy to the sites being fetched: 20 URL counts a minute per visitor.
Ratios, missing counts, and what a chunk here does and doesn't mean.
It depends on the text, which is why the card measures it rather than assuming it. Common English words are usually one token each; longer or rarer words, numbers, names, code and other languages split into several. Checked on 29 September 2026 with OpenAI's o200k_base encoding, Willowdale Equity's home page article ran at 1.33 tokens per word and Verand's home page, heavy with product labels and figures, at 1.52. Paste your own text to see its ratio.
Because Anthropic and Google don't publish tokenizers you can run locally for their current models. Anthropic offers a counting endpoint in its API and describes the result as an estimate, and its documentation says Claude Opus 4.7 and later produce approximately 30 percent more tokens than earlier Claude models for the same text. An OpenAI count relabelled as Claude would be wrong for at least one generation, so this card shows only the counts it can compute exactly.
Because the full page includes everything visible around the article: navigation, header, banners, related links, forms and the footer with its disclosures. On a lean template the article is most of the page; Willowdale's home page is 91 percent article. On a heavy one it can be a small share. If the article share looks implausibly low, read the opening lines on the card: the article region is the page's first article element, then main, and on some templates the first article element is a card or widget rather than the copy.
There's no single answer, and the assistants don't publish one. Some fetch a page when a user asks about it and pass extracted text to the model; search-based answers usually draw on passages retrieved from an index built earlier. What gets kept, trimmed or split differs by product and isn't documented in detail. That's why this card shows both texts and the section sizes rather than claiming to show what any one assistant reads.
No, and nobody outside OpenAI knows exactly how ChatGPT does it. Here the article is cut wherever an H1, H2 or H3 heading starts, each section is counted, and sections under 40 tokens or over 800 are flagged. That's Verand's rule of thumb for passages that can stand on their own, stated as a rule of thumb. It's useful for spotting a section that covers too much or says too little, not as a prediction of any engine's behaviour.
No. In paste mode the text is counted in your browser and no request carrying it is made; the only download is the tokenizer file itself, fetched from jsDelivr's public CDN the first time you count. URL mode is different: our server fetches the public page you name, once, and returns its text to your browser for counting. Nothing is stored either way.
Verand writes from your own experience, gets you found on Google and in AI answers, and tracks where you rank and which answers name you. Every draft passes a compliance gate before anyone can publish it: your industry's rules if you have them, a truth-in-advertising check if you don't.
Researched from the regulators' own text and tested by Verand. Not reviewed by a licensed attorney. Your counsel confirms applicability. Not legal advice. Example shown is illustrative.
The 2019 fund returned 14.2% net to investors, and this offering carries guaranteed returns of 12 to 15% over a five year hold.
Performance advertising can’t promise a return. The claim needs a basis and the required disclosures, and it still can’t be stated as a guarantee.
Willowdale Equity and four competitors · 14 days
Content built to rank in
Google and get cited by
ChatGPT
Perplexity
Gemini
Claude, with every claim checked before it goes live.