Searching a Catalogue by Structure: A Working Chemist's Guide
Share
Searching a Catalogue by Structure:
A Working Chemist's Guide to NorrChemica's Structure Search
How to query our catalogue by drawn structure, SMILES/SMARTS, exact match, identifier, or fingerprint similarity — what each identifier means, when to use it, and why the similarity search is the most useful tool in the panel for anyone planning a compound library.
Most chemical suppliers let you search by name or CAS number. That works only if you already know exactly what you are looking for. A chemist working from a scaffold, a fragment, or a hit compound usually does not — the question is not "do you stock compound X?" but "what do you stock that looks like X?"
NorrChemica's structure search answers that question directly. You draw a molecule or paste a line notation, and the catalogue is queried by chemical structure rather than by text. Every calculation runs inside your own browser; nothing you draw or type is sent to us or stored. This guide explains what each search mode does, what the identifiers mean, and how to get the most out of the feature.
The identifiers, briefly
The panel accepts several ways of specifying a molecule. They are not interchangeable — each encodes something different, and choosing the right one is the difference between a precise hit and a frustrating miss.
What each identifier encodes
- SMILES — a molecule as text A line notation that writes atoms and bonds as a string. Phenylboronic acid is OB(O)c1ccccc1. It is the fastest way to enter a specific molecule without drawing it. A given molecule has one canonical SMILES but many equally valid non-canonical forms — the engine normalises them for you.
- SMARTS — a molecule as a pattern SMILES describes one molecule; SMARTS describes a class of them. It lets you specify atom and bond properties — aromaticity, number of connections, ring membership, charge — rather than fixed atoms. [cX3]B(O)O matches an aromatic carbon bearing a boronic acid, i.e. any arylboronic acid, regardless of what else is on the ring. This is the precise way to express a substructure query.
- InChIKey — a fixed-length structural fingerprint A 27-character hashed form of the IUPAC InChI — phenylboronic acid, for instance, is HXITXNWTGFUOAU-UHFFFAOYSA-N. It is unique to a structure, stereochemistry included. The first 14 characters (here HXITXNWTGFUOAU) encode the connectivity skeleton; the block after the hyphen encodes stereochemistry, isotopes and protonation. Searching the 14-character block finds every stereoisomer and salt of the same skeleton; the full key pins one exact stereochemistry — the rigorous way to ask "is this exactly my compound?"
- CAS Registry Number — a database label A unique identifier assigned by the Chemical Abstracts Service, universal in procurement and on safety documents. Note that it is a registry label, not a structural descriptor: different salts, hydrates or stereoisomers can carry different CAS numbers, and the number itself carries no structure you can compute on. Excellent for confirming a known compound; useless for "find me something similar".
- PubChem CID — a cross-reference key PubChem's integer Compound ID — phenylboronic acid is CID 66827. Convenient when you already have a PubChem record open and want to check whether we carry that exact entry.
How to run a search
The interface has two halves: a structure editor on the left, and a set of search modes on the right. The workflow is the same regardless of mode.
- Provide a structure. Either draw it in the editor, or type a SMILES/SMARTS string in the box beneath it. A string entered in the box overrides whatever is drawn. To change an atom while drawing, click it, type the element symbol — B, N, O, S, and so on — and press Enter; the atom is set directly, with no menu. This is the quickest way to place the boron of a boronic acid: draw a benzene ring, click one ring carbon, type B, press Enter, then add the two hydroxyls. The full periodic table is still there if you prefer to pick from it — select the C (atom) tool, click the atom, and choose the element from the popup.
- Choose a search mode on the right (described below).
- Adjust the similarity threshold if you are in similarity mode — otherwise leave it.
- Press Search. Results appear below as ranked product cards with depictions, identifiers and links. Matched substructures are highlighted on each hit.
Getting these strings out of ChemDraw. You do not have to type a SMILES or InChIKey by hand. Draw the structure in ChemDraw, select it with the lasso, then go to Edit → Copy As and choose the descriptor you want — SMILES, InChI or InChIKey. It lands on the clipboard ready to paste straight into the search box.The search modes
| Mode | Input | What it returns |
|---|---|---|
| Substructure | SMILES / SMARTS | Every product that contains the drawn fragment as a sub-part. SMARTS is recommended for precise queries. |
| Exact match | SMILES | The whole molecule only, stereochemistry included (compared via InChIKey). |
| CAS / InChIKey / CID | Identifier | Direct lookup of a single known compound by its registry or database identifier. |
| Similarity | SMILES / drawing | Every product ranked by how chemically similar it is to your query — with an adjustable cut-off. The most powerful mode; see below. |
The similarity search — and why it matters
Exact and substructure searches answer yes/no questions: is this compound, or this fragment, present? The similarity search answers a graded one — how close is each catalogue compound to my query — and ranks the whole catalogue accordingly. For anyone designing a series rather than ordering a single known reagent, this is the tool that earns its place.
What Tanimoto similarity actually measures
Each molecule is reduced to a Morgan / ECFP4 fingerprint: a bit vector recording which circular atomic environments — every atom together with its neighbours out to two bonds — are present in the structure. Two molecules that share many of the same local environments share many of the same bits.
The Tanimoto coefficient is then simply the number of bits the two fingerprints have in common divided by the number of bits present in either. It runs from 0 (no shared environments) to 1 (identical fingerprints). It is the standard, decades-proven measure of molecular similarity in cheminformatics — not a proprietary heuristic.
Close analogues only
Returns compounds that differ from your query by little more than a substituent. Use when you want drop-in replacements or the nearest matched pairs to a hit.
Same chemotype
Returns recognisable congeners sharing the core. The working range for mapping the available members of a scaffold family before committing to synthesis.
Scaffold hops
Returns distant relatives and bioisosteric cores. Use for idea generation and scaffold-hopping when the obvious analogues are exhausted.
Why this is powerful for library design
Suppose you have a hit and want to build a focused library around it for an SAR study. The classic bottleneck is finding out which analogues are commercially available off the shelf, and which you would have to make. Done by hand, that is a slow trawl through databases and catalogues, one compound at a time.
The similarity search collapses that into a single query. Paste your lead, choose a threshold, and the catalogue returns — ranked — every building block within your chosen radius of chemical similarity. You see immediately which arms of the SAR you can populate from stock and where genuine synthesis is required. Loosen the threshold and the net widens to bioisosteres and scaffold hops you might not have thought to look for; tighten it and you get only the closest matched pairs. It turns "what is near my hit?" from a manual survey into one ranked list.
A worked example
Take phenylboronic acid as a starting point — paste OB(O)c1ccccc1 into the box, or draw it, and select Similarity.
- At a threshold around 0.6, the catalogue returns its arylboronic acids ranked by closeness to the query — the recognisable members of the family.
- Tighten to 0.85 and the list contracts to the nearest analogues only — the close matched pairs.
- Loosen to 0.4 and it widens to take in heteroaryl boronic acids, boronate esters and more distant relatives — a broader survey of what the catalogue offers around that motif.
The same procedure works from any lead structure: it is the fastest way to read off, in one pass, what a catalogue can contribute to a planned series.

Everything runs in your browser
One design decision is worth stating plainly, because it is unusual and it matters for anyone working on undisclosed chemistry: nothing you draw or type leaves your machine.
When you load the page, the catalogue index is sent to your browser once. From then on, every structure perception, substructure match, fingerprint calculation and similarity ranking is computed locally, on your own computer. The query structure is never transmitted to a server and is never stored or logged. For a medicinal chemist probing a confidential target, the structures you search are often the most sensitive information you handle — here they simply never travel.
What the tool is built on
- RDKit The open-source cheminformatics toolkit that is the de facto standard in the field, compiled to run in the browser. It handles structure perception, canonicalisation, substructure matching, InChIKey generation, and the Morgan/ECFP4 fingerprints behind the similarity search.
- Kekule.js The open-source molecular editor providing the in-browser drawing surface, including full periodic-table atom selection.
- Python build pipeline The catalogue index — structures, identifiers and precomputed fingerprints — is generated with a Python/RDKit pipeline from our live product data, so the search always reflects the current catalogue.
Naming the stack is deliberate: the search is built on the same validated, peer-used tools a computational chemist would choose for the job, doing the same calculations they would trust in their own scripts — only with the catalogue already indexed and the work done for you, in your browser.
Try it on your own structures
Draw a molecule, paste a SMILES or SMARTS string, or enter a CAS — and search the NorrChemica catalogue by exact match, substructure, or similarity. It runs entirely in your browser.
Open Structure Search → Request a QuoteNotes
The search runs against NorrChemica's indexed catalogue. Tautomer-distinct forms are treated as different molecules; use substructure search to span tautomers. If a structure is not found, the result panel links directly to a custom-synthesis quote request.
Further reading
For background on boronic acid building blocks and reagent selection in cross-coupling chemistry, see NorrChemica's Lab Journal guide: Choosing Your Boron Source for Suzuki–Miyaura Coupling.
© 2026 SynFinn Discovery Oy / NorrChemica · All rights reserved · Content may not be reproduced without written permission