TL;DR: The Amazon ML Challenge 2026 was an entity resolution problem: 1.7 million reference businesses, about 10 million noisy records from two other sources, and the job of deciding which records belong to which business. The competition ran for three days. I gave myself 24 hours. My pipeline finds about 6 candidates per business with a fine-tuned bi-encoder, scores them with two cross-encoders and a LightGBM stacker, and then picks, for every business, the set of matches with the best expected F0.5. It finished with an overall score of 0.988216 macro F0.5, which put me around rank 200 to 230 on the leaderboard, and 0.992 on a held-out part of the training data. The code is open source.
The Problem
You get business records from three sources. Each record is just an ID, a business name, an address and a country. There are no shared IDs between the sources. Source 1 is a clean, deduplicated reference list. For every Source 1 business you have to list every Source 2 and Source 3 record that refers to the same real business. That can be zero records, one, or many.
The scale is what makes it interesting. Training has 2.2 million Source 1 businesses and 10.3 million Source 2 and Source 3 records. The test set has 1.7 million and about 10 million. Comparing every pair is not an option, so the first half of the problem is finding a small set of candidates per business (blocking), and the second half is deciding which candidates are real.
The metric is F0.5, computed for each Source 1 business and then averaged. F0.5 counts precision about twice as much as recall, so a wrong match hurts more than a missed one. Businesses with no match count too: predicting nothing for them scores 1.0, and predicting anything at all scores 0.
A few more rules shaped the design:
- The test set has a third country, France, which never appears in training. It is 15% of the test businesses.
- The candidate file is reviewed too. A smaller candidate set per business ranks higher, so blocking has to be both tight and high recall.
- No external data at all (no geocoding, no business registries), and every model has to be MIT or Apache 2.0 licensed and at most 8B parameters.
And one rule of my own. The competition ran for three days, but I gave myself 24 hours, start to finish. That changes what “a good idea” means: anything that needs a second day of GPU time is out, every stage has to survive a crash without starting over, and there is no time to guess. Every choice below was measured on a holdout before it went in.
What the Data Looks Like
The data is synthetic: a real-looking business with noise applied to it several times. Here are two true match groups from the training data:
S1 Payne Enterprises 3315 Fremont Street, Peoria, IL
S2 Payne Énterprises 3315 FREMONT ST, PEORIA, IL
S2 PAYNE-ENRTPRMISES 3315 FREMONT SAINT, PEORIA, IL
S3 Payne Etrepndiels 3315 Fremont St, Peoria, Illinois
S3 Payne Enterprises LLC Fremont St, Peoria, Illinois
S1 Raj Investments LLP 6(29), C.I.T. Colony, 2Nd Main Road Mylapore, Chennai, Tamil Nadu
S2 ராஜ் இன்வெஸ்ட்மெண்ட்ஸ் எல்எல்பி 6(29), C.I.T. COLONY, 2ND MAIN ROAD MYLAPORE, CHENNAI, Tamil Nadu
S3 Raj Investments எல்எல்பி 6(29), C.i.t. Colony, 2Nd Main Road Mylapore, Chennai, TN
S3 ராஜ் இன்வெஸ்ட்மெண்ட்ஸ் எல்எல்பி 6(29), C.i.t. Colony, 2Nd Main Road Mylapore, Chennai, தமிழ்நாடு
Typos, injected accents, “St” expanded to “Saint”, a
dropped house number, a legal form added, and for India, names and state names
written in Tamil, Devanagari, Kannada and other scripts (28% of Indian Source 2
names). On top of that: junk around names (-- ,
<< , | www.x.com), legal forms moved to the
front (“LLC Moncada Léarning Center”), names written as a domain
(maurewilliamscolombier.com), reordered address parts,
null fillers, empty addresses, and records with a completely
different trade name at the same address.
The hard part is not the noise, though. It is the distractors. About 26% of Source 2 and Source 3 records match nothing, and they are not random. They are near copies of a real business: a house number a few units off (1024 vs 1027), one real word swapped (First vs Seventh, Private vs Public), or a word like “& Sons” or “Group” added. In training, a record that adds “& Sons”, “& Associates”, “Group” or “Enterprises” to the business name at the same address is a distractor 98 to 100% of the time, while adding “Services” or “Center” is ordinary noise. And about half of all business names are shared by at least two different businesses in the same country, so a name alone is often not enough.
The Most Useful Fact
The training ground truth has 7,638,365 matched record IDs, and every single one is unique. Each Source 2 or Source 3 record belongs to at most one business.
That flips the problem around. Instead of only asking “which records match this business?”, you can ask “which business owns this record, if any?”. The pipeline uses this in three places: search runs in both directions (business to records and record to businesses), the models get features about how strongly other businesses compete for a record, and at the end each record is given to one business only.
Validation
Everything was measured on a split of the training businesses, and the split matters more than usual because the pipeline stacks models on top of models:
| Split | Share | Used for |
|---|---|---|
| A | 80% | training the bi-encoder and cross-encoders |
| B | 10% | training the pruning model and the stacker |
| C1 | 8% | calibration and choosing the final sets |
| C2 | 2% | untouched, only for checking |
The stacker has to learn from neural scores on businesses the neural models never saw, otherwise it learns to trust scores that are too good. Retrieval for B and C searches the full training pool, distractors included, so the holdout sees the same kind of search the test set does.
France has no labels, so I used a stand-in: train on US only and score India, and the other way round. Choices that helped US and India but hurt this cross-country transfer were left out.
The Pipeline
1. Normalization
Every record gets a cleaned version next to the raw one: Unicode NFKC, lowercase,
accents stripped, non-Latin scripts transliterated with anyascii,
junk prefixes and suffixes removed, null-style fillers dropped,
and repeated words collapsed. Legal forms (LLC, Inc, Pvt Ltd, SARL, SAS and so
on) are found anywhere in the name and moved to their own field, which leaves a
“core name”. Aliases (“X dba Y”, “t/a”,
“formerly”) and domain stems get their own fields too.
Transliteration alone is not enough. लिमिटेड comes out as something
like limitedd and Tamil Nadu comes out as tmilnatu. So
instead of writing a dictionary by hand, the pipeline mines one from the
training pairs: it lines up tokens between records that are known to
match and keeps the mappings that show up consistently. That gives a 555-entry
native-to-Latin word map, state aliases (tmilnatu and
tn to tamil nadu) and street fixes (aveune to avenue).
Some real entries from the mined maps:
| After transliteration | Mined as | Kind |
|---|---|---|
ailailpi | llp | native-script word |
aimtrpraij'ij' | enterprises | native-script word |
aiksports | exports | native-script word |
ainrji | energy | native-script word |
bombay, calcutta | mumbai, kolkata | city alias |
flor, flt | floor, flat | address typo |
aveune, ausin | avenue, austin | address typo |
All mined maps are keyed by the country string, so nothing is hard-coded to US or India. A short hand-written list covers French street abbreviations and maps French business words to English ones (“& Fils” to “& Sons”, “Groupe” to “Group”), so the distractor patterns the models learned in English also apply to French names.
2. Candidate Generation
Retrieval uses intfloat/multilingual-e5-small fine-tuned as a
bi-encoder on split A. Round one uses in-batch negatives with batches of 1,024
pairs. Each batch holds one country and unique businesses, so no in-batch
negative is ever a true match. Round two continues from round one with hard
negatives mined from its own search results, and trains three views of each
record: the full record (60% of batches), the name only (20%) and the address
only (20%). The views matter because some records have no address, and some have
a totally different trade name at the same address.
Every record is encoded in four views (round-two full, name and address, plus
round one), and exact kNN runs on the GPU inside each country. It runs in both
directions: each business looks for its top records, and each record looks for
its top businesses. Because each record has one owner, the reverse direction is a
strong filter. Two exact keys catch what embeddings miss: an acronym of the core
name (for records like GAS) and a domain stem equal to the core name
without spaces.
The union keeps 99.92% of true pairs but at 58 candidates per business, which is far too many. A small LightGBM model trained on split B cuts it down using only cheap signals: the cosine and both ranks of every view, four fast rapidfuzz scores, and rank and gap context. Pairs scoring at least 0.002 are kept, at most 15 per business:
The output of this model is the candidate file, and it is exactly the set the rest of the pipeline scores. For comparison, the same e5 model without fine-tuning needs 60 candidates per business to reach 98.7%.
3. Scoring Pairs
Each candidate pair gets a wide set of features:
- Names: rapidfuzz ratio, partial ratio, token set, token sort and Jaro-Winkler on the cleaned name, the core name and the raw name; alias and domain matches; legal form equal or missing; idf-weighted token overlap.
- A word-swap signal: among the record’s tokens with no fuzzy partner in the business name, take the smallest idf. If the unmatched word is common (“Seventh”), it was probably swapped, which points to a distractor. If it is rare, it is probably just a typo.
- Addresses: the same string scores, idf overlap, first house number equal, number overlap, dropped leading digits (2880 vs 880), the size of a house number difference, phone digits, and empty-address flags.
- Retrieval: the four bi-encoder cosines, their ranks in both directions, and the pruning score.
- Context: the rank, gap and margin of the main scores within the business’s candidates and within the record’s candidates.
On top of that there are two cross-encoders, both
microsoft/mdeberta-v3-base (278M parameters). One reads the cleaned
text and one reads the raw text, since the raw version still has signals that
cleaning throws away. Each is trained for one epoch on 3 million split-A
candidate pairs, 40% of them positive, with the negatives coming from the
retrieval lists so they are hard. One practical note: recent
transformers versions load mdeberta-v3 in fp16, and training it that
way gives NaN losses. Loading it in fp32 and training with bf16 autocast works.
So what does the stage-1 stacker actually lean on? Grouping its 131 features by where they come from:
4. Stacking With Group Features
A LightGBM stacker, trained 4-fold on split B, combines everything. Then a second stage adds group features, which come from one observation: true records agree with each other, and a distractor has a change nobody else shares. If four records of a business all say house number 1024 and a fifth says 1027, the fifth one is suspicious, even if 1027 is close to what the business itself says.
So for every candidate, stage 2 looks at the business’s other confident records from stage 1 (out-of-fold predictions, so there is no leakage) and measures how similar the candidate is to them:
strong = p.filter((pl.col("p") >= 0.5) & pl.col("is_arg")).select("i1", pl.col("i2").alias("j"))
pairs = p.select("i1", "i2").join(strong, on="i1").filter(pl.col("i2") != pl.col("j"))
...
g = pairs.group_by("i1", "i2").agg(
pl.col("rc").max().alias("g_rc_max"), pl.col("rc").mean().alias("g_rc_mean"),
pl.col("rn").max().alias("g_rn_max"), pl.col("rn").mean().alias("g_rn_mean"),
pl.col("ra").max().alias("g_ra_max"), pl.col("ra").mean().alias("g_ra_mean"),
pl.col("rnum").filter(pl.col("rnum_ok")).mean().alias("g_num_support"),
...
)
rc, rn and ra are the embedding, name and
address similarity between the candidate and each confident sibling, and
g_num_support is the share of siblings with the same house number.
There is also a feature for whether the business’s own house number agrees
with its records, which tells the model when the reference record is the noisy
one. On the record side, stage 2 adds the probability of the best
other business for this record and the margin to it.
5. Picking the Final Sets
The stacker gives a probability for each pair. Turning those into lists takes three steps. First, one-to-one: each record is kept only for the business where it scores highest.
def one_to_one(df):
return df.sort(["i2", "p", "i1"], descending=[False, True, False]).unique("i2", keep="first", maintain_order=True)
Second, the probabilities are calibrated with isotonic regression on C1. Third, instead of one global threshold, each business gets the set that maximizes its expected F0.5. Sort its candidates by probability q. For the top-k set, the expected number of true matches in the set is roughly the sum of their q, and the expected number of true matches overall is the sum of all q. Plug those into the F0.5 formula and try every k. The empty set is also a choice: it scores 1.0 only if there is no true match at all, which has probability ∏(1 − q).
d = df.sort(["i1", "q"], descending=[False, True]).with_columns(
pl.col("q").cum_sum().over("i1").alias("cs"),
pl.col("q").cum_count().over("i1").alias("k"),
pl.col("q").sum().over("i1").alias("tot"),
(1 - pl.col("q")).log().sum().over("i1").exp().alias("p_empty"),
)
d = d.with_columns((1.25 * pl.col("cs") / (pl.col("k") + 0.25 * pl.col("tot"))).alias("ef"))
This handles businesses with no match without any special case: when every candidate is weak, the empty set wins. There is also an exact version that computes the full Poisson-binomial expectation on the GPU. Which one to use is decided on C1, and they differ by less than 0.0001.
6. Two Runs, and France
The whole pipeline (bi-encoders, cross-encoders, pruning model, stackers) is trained twice with different seeds. For US and India, the stage-2 logits of the two runs are averaged. A pair that one run did not keep as a candidate counts as probability 0.0001 in that run, and the candidate file is the union of both runs’ candidates, since that is the set the average scores.
France is handled differently. Two runs trained the same way disagree on 8% of French output rows, against 0.8% for US and India. The pairs only one run accepts sit right around the point where adding a pair stops paying off under F0.5, and there are no labels to check them against. So for France the final file keeps only the pairs both runs accept:
pp = p1.join(p2, on=["source1_entity_id", "matched_entity_ids"]) if a.mode == "intersect" else pl.concat([p1, p2]).unique()
Results
| Macro F0.5 | |
|---|---|
| Final overall score (all test countries) | 0.988216 |
| Holdout C1 (US and India) | 0.99254 |
| Holdout C2, untouched (US and India) | 0.99219 |
| US, C1 / C2 | 0.9917 / 0.9914 |
| India, C1 / C2 | 0.9937 / 0.9934 |
On the holdout, pair-level precision is 0.9993 and recall is 0.977. The pipeline is very careful: it would rather miss a record than claim a wrong one, which is what F0.5 asks for. On the test set, blocking produces about 6.5 candidates per business per run.
The two cross-country checks (train on one country, score the other) come out at 0.986 for US to India and 0.982 for India to US. The fine-tuned bi-encoder features transfer especially well.
Where It Still Fails
Legal-form distractors. The most common wrong merge is a distractor that is the business name plus “Inc” or “LLC” at the same address. True records get an added legal form just as often, and sibling records share it about equally often in both cases (4.2% vs 4.0%). I could not find a signal that separates them.
Records with no address. True pairs with an address are recovered 99.8% of the time. Pairs where the record has no address, only 64%. Most of the misses have a name shared by several businesses in the same country, and the best weak signal I found (the right business has fewer records from the same source) is right only 39% of the time. With F0.5, predicting nothing is the correct call there.
France. The holdout says about 0.992 for US and India, and the accepted US and India test pairs look just like the holdout ones (same rates of word swaps, number changes, legal-form changes and empty addresses, same mean probability, same 3.4 matches per business). That puts France at roughly 0.96, and it is where most of the leaderboard gap comes from. French addresses end with a long region name that records drop 65% of the time or replace with a department (“Nord”, “Gironde”), and 13% of French businesses share an exact address with another one, against 4 to 6% in training. The French errors are mostly pairs the model is confidently wrong about, not threshold choices.
If I had more time, these are the things I would try next:
- Inject French-style noise into Indian data, train on US, and use that as a closer stand-in for France.
- Features that combine how many businesses share an address with how their names differ.
- A dedicated model for French distractor patterns.
Where the 24 Hours Go
One full run of the pipeline, from raw TSVs to the two output files, takes just under 8 hours on one GPU. The final output averages two of them, so that is most of a day of GPU time on its own:
That is why every stage writes its output to disk and skips work that is already done. If a job dies halfway, rerunning the same command picks up from the last finished stage instead of costing hours.
Try It
You will need the competition dataset, a Linux machine with one GPU with at least 40 GB of memory, about 100 GB of RAM and 150 GB of free disk.
git clone https://github.com/ArneshBanerjee/amazon-ml-challenge-2026-entity-resolution
cd amazon-ml-challenge-2026-entity-resolution
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements.txt \
--extra-index-url https://download.pytorch.org/whl/cu126 --index-strategy unsafe-best-match
# two runs, averaged (the final output)
PY=$PWD/.venv/bin/python bash src/run_ensemble.sh /path/to/dataset /path/to/output
One run takes 8 to 10 hours on one GPU, most of it transformer training and
inference. Every stage writes its output to disk and skips work that is already
done, so if the machine goes down you can run the same command again and it
continues where it stopped. run_all.sh does a single run in about
half the time.
Credits
The challenge and dataset are from the Amazon ML Challenge 2026. The models are
intfloat/multilingual-e5-small and
microsoft/mdeberta-v3-base,
both MIT licensed, along with LightGBM, polars, rapidfuzz, PyTorch and
transformers. No external data was used.
Code:
github.com/ArneshBanerjee/amazon-ml-challenge-2026-entity-resolution
(MIT)
Author:
arneshbanerjee.dev
· Contact:
arneshbanerjee24 [at] gmail [dot] com
Loading comments…