An engineering works selling into export markets can have a flawless catalogue and still be half invisible. A missing page does not crash, slow down or send an error — it drops out quietly, and only someone who counts will notice.
This text is about what happens before rankings can be discussed: whether a page enters the search engine's collection. The examples come from diversified manufacturing, where items run into thousands and the market is the whole country plus export. The companies are invented.
The fault that shows in no metric
Take a subcontract engineering works in Ostrobothnia that rebuilt its site in the spring. The whole range went live at once, links work, pages load in a second. Six months on, sales asks why enquiries are down. Analytics shows no drop, because there is none: nobody ever came to those pages.
That is the nastiest feature of an indexing problem. A broken form is noticed in a day, a crashed server in minutes, but a page outside the collection behaves like one nobody needs. Both read as zero, and the difference shows only when published addresses are compared with findable ones.
In manufacturing that gap is usually large, for structural reasons. Dozens of product families, each with sizes, materials and a datasheet, in four languages because Germany and Sweden buy in their own. Ten families swell into thousands of addresses without one editorial decision, and no system asks whether all belong.
Discovery, fetch and admission to the collection
The commonest misunderstanding is that indexing is one event. There are three consecutive stages, each able to stall for its own reasons. Kept apart, the question becomes answerable: which step breaks?
- Stage one: discovery. The search engine learns the address exists, by four routes: links in your navigation and body text, a sitemap row, an outside link, or a notification. Without one of them the page exists for nobody but you.
- Stage two: fetch. The robot makes an actual request to your server. How many requests your domain gets in a given time is the crawl budget — an estimate from response speed, past errors and the domain's judged value.
- Stage three: admission. The fetched page is assessed and either kept or left out, on how distinctive the content is, how many near-identical versions exist, and whether its technical signals agree.
The gap between stages can be long. A row in the sitemap since September may be unfetched in December, while another has been fetched cleanly three times and still does not appear. They need different remedies but, unmeasured, look alike: the page is not on Google.
Crawl budget cannot be reached directly. You influence it by speeding up responses, cutting pointless addresses and keeping the robot from spending ten visits on one piece of content. On most Finnish business sites the budget is not the bottleneck — where it goes is.
Why a page stays outside the collection
Large industrial sites show the same causes again and again. They are grouped by the stage at which the chain breaks, since that decides who fixes it: server side, CMS administrator or content team.
| Cause | Breaks at | How to spot it | Who fixes it |
|---|---|---|---|
| Item reachable only via the search box | Discovery | Opens at a direct address, but no browsable path leads there | CMS: browsable group pages |
| Datasheets exist only as attachments | Discovery | Technical data lives in files nothing links to | Content team |
| Size and material choices are parameters | Admission | One item answers at dozens of parameter addresses | Development: canonical |
| Language versions point at each other wrongly | Admission | The German page names English as canonical | Development |
| Staging environment open to the web | Admission | The same content answers from two domains | Server side |
| The robots file is used for hiding | Fetch | Still shows in search although blocked in robots.txt | Server side |
| Response time stretches in the catalogue | Fetch | Filter views answer in seconds, items quickly | Development: caching |
Two rows deserve opening up. A robots file is not a hiding tool: block an address there and the robot stops reading the page, so it never sees the instruction inside to stay out. The address hangs on in the results without a description. To remove a page, leave access open.
The second is the canonical reference, which many take for an order. It is a suggestion: the engine picks the version to keep, on a multilingual export site often not yours. The typical case is a product family whose Finnish, Swedish and German pages follow one template so faithfully that the engine keeps one.
A thousand a day, ten thousand at once
The indexing part of the Semalt panel is three things: submitting addresses, processing sitemap files, and a log of what the robots did next. It turns useful once you read the stated limits, which say what fits where.
Submitting and tracking addresses
Addresses you want known now, listed by hand or exported from a system.
- The daily quota is a thousand addresses. The limit applies to the account, not one site. A parent company and three subsidiaries on one account share it, and the order is yours.
- A batch holds ten thousand rows. Batch size and daily quota are different measures. A large batch saves your time but speeds nothing up: ten thousand rows unwind over ten days.
- Three counters show the state. Submitted, found and failed update during the run, so failures show before the batch is done.
- In ordinary weeks the quota is ample. Ten new items a week and two updated group pages come nowhere near it. Scarcity arises in three cases: platform change, domain move, first full run.
The quota feels like a restriction, but it forces the order to be decided. An unsorted queue is processed as it arose, in an order set by the CMS rather than by sales. When a thousand fit in a day, you must ask which thousand.
Address lists can also go to the Stream assistant in My SEO, which takes them in batches. The quota is the same whichever way they arrive.
Sitemap work: three levels, a thousand files, two in parallel
The second route starts from sitemaps. It suits a case where the division exists already: product group, language, content type. A large site produces these files automatically, which is both the route's strength and its pitfall.
Sitemap processing as a job
A full run when the structure already exists in a file tree.
- Upload or address. The file can be sent from your machine or named by its address. The first suits trial runs, before a new structure is public.
- Nesting unwinds to three levels. An index file may point to another, and the trail is followed three layers down. Deeper structures are worth flattening anyway, since nobody can read them.
- A thousand files in one job. That holds four languages, thirty product groups and separate files for documents and news without meeting the limit.
- Two jobs at once, twenty waiting. Parallelism is capped at two, the queue holds twenty. For an agency with several clients, an urgent run belongs ahead of routine checks.
Either route, entries land in the same log and the quota drains from the same vessel. The difference is who chose the rows: you assemble the list, the CMS assembles the sitemap, and it knows nothing about what sells now.
A sitemap is not a stock list but a claim: every row says this address deserves its place. When a third of them are redirects, discontinued items and addresses marked to be left out, the claim loses value — and the loss reaches the good rows too. Cleaning before submission is a condition of the result.
A file per language and product group
Clustered errors show which part of the catalogue they hit; one large file says only that something is wrong somewhere.
- An index file holds the pieces together
- A separate file for documents
Final addresses only
A row must answer directly, be final, and be the version you named to keep.
- No intermediate steps of redirect chains
- No variants separated by parameters
A change date you can trust
If the overnight run restamps every row, the field means nothing. Set it when text or price actually changed.
- An overnight run is not a content change
- A wrong date wastes fetches
Out with what does not belong in search
Rows the engine rejects anyway eat the file's credibility and return nothing.
- Blocked and excluded addresses
- Logged-in user views
- Filter views and internal search
So the order is: run through, fix, then spend the quota. The first run is an inventory — rows held, rows that failed, where errors cluster — and only a cleaned list makes the quota productive.
IndexNow turns the initiative around
By default the search engine decides when it comes, and on a rarely updated domain the interval stretches to weeks, a new item low in the structure waiting longest. IndexNow reverses this: the domain sends its own message on a change, accepted by both GoogleBot and BingBot.
The value of a notice drops if it becomes routine. The right moment is when the page's substance is genuinely different from yesterday; a footer year changing is not. Two robots come on one notice, but the quota is the same whatever route the address took. How Semalt's indexing tool fits with the other sections is clearest when the search console figures come from the same account.
The benefit is not speed but a sharper question: the page was fetched and not kept, so what is wrong with it?
A launch with a date on it
An engineering trade-fair launch is timed to the hour. The page must be findable that day, not weeks later.
- Submit at the moment of publication
- Prepare the address in advance
The language version nobody watches
Of four languages, one always has figures nobody reads. The backlog piles up there.
- Name an owner for the language
- Count its indexing rate separately
Dealer details change
A new partner in Kuopio or a changed service point in Rovaniemi alters the page genuinely.
- Submit on a real change
- Not on reformatted contact details
An address just opened
A block lifted, an exclusion removed, or a group page with no link to it yet.
- Straight after the fix
- Together with internal linking
Time, status code and error text per address
The real benefit comes not from sending but from what you see afterwards. Each address gathers a record of when the robot visited, which status code it got and what error text came back. An opinion that Google dislikes our pages does not survive one log row.
- Submitted, never fetched. The counter shows the submission, the log row is missing. Check the robots file, server load, and any throttling after earlier errors.
- Fetched during a maintenance window. A five-hundred response is not one page's problem. It affects how cautiously the robot treats the whole domain for weeks.
- Fetched, but nothing was there. Four-hundred responses expose a stale sitemap row. Take it out of the file and do not submit that address again.
- Fetched, but redirected onward. You submitted a waypoint, not the destination. Put the final address into the sitemap and the list.
- Fetched and rejected. The technical side is in order, so three options remain: content too thin, the canonical pointing elsewhere, or no internal link to the page.
Tie the reading pace to the phase. During a large run twice a week is right, so error clusters get fixed before they eat days of quota. Afterwards a monthly check on new items and the busiest groups suffices.
One technical point: the per-address visit log sits in the same account as the search console and rank-tracking views, so indexing and visibility compare without a separate tool. When a group page has been findable three weeks and no impressions arrive, the problem is content, not indexing.
The points most often asked
We have four language versions. Does each have to be handled separately?
Yes. Languages are independent addresses and all eat from the same daily quota. Give each its own file under one index file, using two of the three permitted levels. Sales decides the order, and the language with the highest enquiry value goes first. Hreflang markings belong in the page template.
The catalogue updates overnight from the ERP. Should we submit the whole catalogue weekly?
No. The overnight run touches thousands of rows, but perhaps thirty change in substance. Submit everything and the quota goes on items where only the stock balance moved. Build a comparison into the update run that picks the genuinely changed rows; it pays for itself in a month.
Our location pages do not reach search. Is the fault in the submission?
Check the log first: if the pages have been fetched, the problem is content. When the only difference between two pages is the town name, they are read as one and one is kept. Keep the locations with something of their own — a dealer, a service point, a delivery time — and merge the rest, since twenty distinctive area pages beat two hundred formulaic ones.
How long should we wait before a missing page is really a problem?
It depends on the stage. A week since submission with no visit row in the log is a discovery problem, worth solving at once. If the fetch happened but the page is not in search after two weeks, look at content and canonical references. Catalogue-wide indexing rate wants checking monthly.
Can two sitemap jobs run at once for different domains?
They can: two parallel jobs are allowed and the queue holds twenty, enough for most groups and agencies provided somebody decides the order. Give priority to the domain with a change under way — a platform switch, a new product family, a domain move. A sitemap run in the panel continues in the queue by itself, so nothing needs watching once started.
What a large catalogue costs in the calendar
Finally a worked example. An imaginary engineering group has forty thousand addresses across four languages. The quota is a thousand a day, so the rough answer would be forty days. It is wrong, because it assumes every address deserves a place in the queue.
| Address set | Count | Belongs in the queue | Action |
|---|---|---|---|
| Product pages in four languages | 16,800 | Yes, in stages | Finnish and English first, then German and Swedish |
| Size and material choices as parameters | 11,500 | Not at all | Canonical reference to the base item |
| Location and dealer pages | 3,800 | Only after pruning | Distinctive ones stay, rest merged |
| Technical document pages | 3,600 | Yes, second wave | Their own sitemap file |
| News and reference archive | 2,400 | Yes, background | No deadline |
| Discontinued item addresses | 1,900 | Not at all | Redirect to the successor or a gone response |
Two sets drop out at once: parameter variants and discontinued items, 13,400 addresses solved by structure rather than submitting. That leaves 26,600 submittable addresses, or 27 days of quota — a different promise to the board than forty.
The first wave matters more. The Finnish and English pages of 4,200 items plus 320 group pages make 8,720 addresses, or nine days — after which the part of the catalogue producing most enquiries is findable. German and Swedish follow, documents and archive behind them.
Indexing stays underrated because it raises no alarm. A page that is not findable looks in the statistics like one nobody asks for, so half a year goes on polishing texts while the obstacle sits a step earlier. The quota, two parallel jobs and three sitemap levels are limits to live with — but they are numbers you can calculate with, and only a calculated schedule can be promised. More technical reviews appear on the blog.
Start by measuring rather than sending: open the Semalt panel and run your catalogue's sitemaps through it, and the first thing you see is the gap nobody has counted. It is usually larger than you guessed, and the only proper starting point for everything else.