The short version
- The PDF catalogue is the company's richest information asset and its most invisible.
- A PDF can be indexed, but it competes poorly, does not update, and takes the visitor off your site.
- The rule: HTML to be found, PDF to be downloaded. You need both, in different roles.
- Every part number is a query. A 400-item catalogue is a 400-page opportunity.
- Anyone not publishing extractable data cannot be cited by a conversational assistant.
Almost every manufacturer has the same archive: a general PDF catalogue, a dozen datasheets, some dimensioned drawings, a restricted price list. Carefully maintained, accurate documents — and to a search engine, almost non-existent.
The paradox is that this material contains exactly what customers search for: part numbers, dimensions, tolerances, materials, operating conditions. The company's most valuable content sits locked in a format that cannot compete.
This article covers why that happens, what changes when the same data is published in HTML, and how to do it without rebuilding the catalogue from scratch.
What actually happens to a PDF
The common belief is that PDFs are not indexed. Not true: they are. The problem is different, and subtler.
| Aspect | HTML page | |
|---|---|---|
| Indexing | Possible, but only if the text is text | Full |
| Scans and images | Unreadable without OCR: content lost | Not applicable |
| Title and description in results | Often the file name | Under your control |
| Internal links | Absent: the PDF is a dead end | Lead to other pages and to contact |
| Updating | Regenerate and re-upload everything | Edit the single value |
| Reading on a phone | Constant zooming, often abandoned | Natural |
| Measurability | Almost none | Impressions, clicks, position per page |
A catalogue exported as images or scanned. The text does not exist as text: no engine can read it, no assistant can cite it, no part-number search will ever find it. It is the format in which most companies keep their information assets.
The operating rule
This is not about eliminating PDFs: in B2B they are needed, because the engineer attaches them to the file, purchasing archives them and sales emails them. It is about giving each format the right role.
HTML to be found. The data lives on the page: readable, updatable, linkable, measurable.
PDF to be downloaded. The same content, formatted for print and archiving, offered from the page rather than instead of it.
This simple swap of priorities produces three effects: searches on specific data find a page instead of a file; the download becomes a trackable action and therefore an interest signal; and the visitor stays on the site, where links to related pages and to contact exist.
Every part number is a query
This is the point that changes the scale of the argument. In B2B, someone searching a part number knows exactly what they need: they are replacing a component, looking for an equivalent, or checking compatibility. It is the purest commercial intent there is, with almost no competition.
A 400-item catalogue is therefore a 400-page opportunity. Which, put that way, is alarming — and is exactly why the project never starts.
The solution is not writing them by hand but treating them as data. The minimum structure of a part page is always the same: designation and synonyms, dimensional data, material and treatment, typical applications, related or alternative codes, availability and minimum batch, and a link to the corresponding capability page. Generated from a table, it is days of technical work rather than months.
Publishing four hundred pages in one day on an industrial domain with modest authority produces slow indexing and uncertain evaluation. Start with the fifty codes that generate the most enquiries and proceed in batches, confirming each enters the index before loading the next.
Building a part page without writing it by hand
The fear that blocks this project is always the same: "we cannot write four hundred pages." Correct, and they should not be written. They should be generated from data the company already holds, with editorial work concentrated on three points only.
The source. In most cases a spreadsheet or an ERP extract already exists with code, description, dimensions, material and availability. That table is the base. If it does not exist, building it is an investment that pays elsewhere too: quotations, price lists, customer integrations.
The page template. A single structure into which the table fields flow automatically: title with the designation, data table, applications section, related codes, link to the process. This part is development work, done once.
The text that is not generated. Here lies the difference between an online catalogue and one that works. For each family — not each code — you need a paragraph written by someone who knows the product: what it is for, in what context, which problem it solves, when another choice is better. Twenty families means twenty paragraphs, a day's work with the engineering office. That text, inherited by every page in the family, is what makes them something other than a list of numbers.
Finally, a decision to take deliberately: which codes not to publish. Discontinued items, bespoke work, single-customer references. Putting everything online because "it is in the ERP anyway" fills the index with pages that compete for nothing and dilutes the rest. A public catalogue is a selection, not a dump.
The duplication risk, which here is real
Four hundred pages differing only by a number are, to a search engine, one page. It is the most serious technical problem in this kind of project and it has to be addressed before generating, not after.
What makes a part page sufficiently distinct: genuinely different numeric data rather than just an identifier; applications specific to that code; related codes that differ page to page; and, where it exists, its own image or drawing.
What is not enough: changing the title, repeating the same paragraph with the code swapped, adding generic text. Those are the three most common solutions and none of them works.
When variants are genuinely minimal — the same part in five lengths — the correct choice is a single family page with a variants table, not five near-identical pages. Fewer pages, each with a reason to exist, always beats the reverse.
Verifying they are actually indexed
With volume generation the risk is not producing too little: it is producing a lot and never noticing that half of it never entered the index. Verification runs in three steps, in this order.
- Are the bots coming?The crawl log by section shows whether crawlers are actually scanning the new area. If they are not, it is a discovery problem: missing internal links, or an un-updated sitemap.
- Do they enter the index?Submitted versus indexed URLs. Below 85% on a technical catalogue means the pages resemble each other too closely: the data that should distinguish them is missing.
- Do they collect impressions?Indexed with no impressions means they compete for nothing. On a catalogue, that signals the chosen codes are not the ones the market searches for.
Three outcomes, three different actions. Without this sequence the typical reaction is "let's add more pages", which in cases one and three makes things worse.
Drawings, images and the visual half of a catalogue
A chapter almost nobody treats as SEO, and which in technical B2B matters more than elsewhere: engineers often search by shape, not only by word.
Three habits change the outcome without requiring new photography. File naming: an image called IMG_4471.jpg says nothing, one called flanged-bronze-bush-b12-section.jpg describes its content. Alternative text, which should describe the part and the view rather than repeat the company name. And a visible caption, which on a technical drawing also serves the reader: which view, which scale, which dimensions are indicative.
There is also a substantive point: publishing dimensioned drawings alarms many companies, who fear handing information to competitors. Worth distinguishing. Envelope dimensions and interfaces let the customer check compatibility and are in practice already known to the market; the production process, machining parameters and costs should not be published and nobody expects them to be. Confidentiality lies in how you produce, not in what you sell.
A last practical note: heavy images are the main cause of slowness on technical catalogues, because they are served at original CAD or camera dimensions. Resizing them to their actual display size and compressing them is half a day's work that, across hundreds of pages, noticeably changes the experience for someone consulting from a workshop on a mediocre connection.
The PDF stays, but changes position
Some practical guidance for documents that remain as PDFs, because they can still work better than they do today.
The text must be text. A PDF exported from CAD or the ERP keeps selectable text; a scan does not. If you cannot select a word in your own catalogue, that document is blind to every system.
A readable file name. datasheet-flanged-bronze-bush-b12.pdf rather than doc_final_v3.pdf. It appears in results and in downloads.
A title in the document properties. Many PDFs carry the source file path as their title, and that is what ends up in search results.
Reachable from a page. A PDF linked only from the sitemap is an orphan. It should be downloadable from the HTML page covering the same subject.
One location only. The same catalogue uploaded into four folders creates four competing URLs for one document.
Technical documents are what assistants cite
There is a recent reason to accelerate this work, concerning a channel that did not exist two years ago.
When an engineer asks a conversational assistant for a supplier with particular characteristics — a tolerance, a material, a certification, a lead time — the answer is built by citing sources. A system can only cite what it can read and extract with confidence: structured data, as text, on a page.
Which means a scanned PDF catalogue is not merely less visible on Google. It is entirely absent from a growing channel, and will remain so when that channel matters considerably more than it does today.
If measuring it is worthwhile, the AI Analytics module shows which questions in your category produce answers and who gets cited. In industrial niches it is common for no actual manufacturer to be cited at all: answers fall back on generic portals. That empty space gets occupied by publishing data, not adjectives.
Check whether your catalogue is readable
Try selecting a word in your PDF. Then check the indexing log to see whether bots visit the technical section of your site. Two checks, ten minutes, free.
Sign in to Semalt See Fast IndexingFrequently asked questions
Are PDFs indexed?
Yes, if they contain real text. But they compete worse than HTML pages, do not update easily, and lead nowhere.
Should I remove PDF catalogues?
No. Pair them: HTML to be found, a PDF downloadable from the page for the customer's archive.
How many part pages should I create?
Start with the fifty codes that generate the most enquiries, verify indexing and impressions, then continue in batches.
What if variants are nearly identical?
A family page with a variants table. Five pages differing by one digit are duplicates.
How do I know it is working?
Three checks in sequence: bots visiting, pages entering the index, impressions arriving. Each outcome points to a different action.
Conclusion
In most industrial companies the content that would work already exists: it is in the catalogue, the datasheets, the drawings. The material is not missing, the format is. Moving that same data onto the page is one of the few SEO interventions that requires inventing nothing — it requires making readable what has already been written.
Today's check costs ten minutes: open your catalogue and try to select a word. Then connect the domain and see whether bots visit the technical section. The two answers together tell you where to start.
If you would like us to review your catalogue and estimate how many pages are worth generating, get in touch: the initial audit is free and delivered within 24 hours.