Blog

Amazon ML Challenge 2026 in a Self‑Imposed 24 Hours: 0.988216 F0.5

TL;DR: The Amazon ML Challenge 2026 was an entity resolution problem: 1.7 million reference businesses, about 10 million noisy records from two other sources, and the job of deciding which records belong to which business. The competition ran for three days. I gave myself 24 hours. My pipeline finds about 6 candidates per business with a fine-tuned bi-encoder, scores them with two cross-encoders and a LightGBM stacker, and then picks, for every business, the set of matches with the best expected F0.5. It finished with an overall score of 0.988216 macro F0.5, which put me around rank 200 to 230 on the leaderboard, and 0.992 on a held-out part of the training data. The code is open source.

The Problem

You get business records from three sources. Each record is just an ID, a business name, an address and a country. There are no shared IDs between the sources. Source 1 is a clean, deduplicated reference list. For every Source 1 business you have to list every Source 2 and Source 3 record that refers to the same real business. That can be zero records, one, or many.

The scale is what makes it interesting. Training has 2.2 million Source 1 businesses and 10.3 million Source 2 and Source 3 records. The test set has 1.7 million and about 10 million. Comparing every pair is not an option, so the first half of the problem is finding a small set of candidates per business (blocking), and the second half is deciding which candidates are real.

The metric is F0.5, computed for each Source 1 business and then averaged. F0.5 counts precision about twice as much as recall, so a wrong match hurts more than a missed one. Businesses with no match count too: predicting nothing for them scores 1.0, and predicting anything at all scores 0.

Number of true matches per Source 1 business in the training data 0 200k 400k 600k 0 1 2 3 4 5 6 7 8 9 10 11 5.6% peak: 3 (24%) true matches per business
How many Source 2 and Source 3 records each training business really has. Most have two to five, and 5.6% have none at all, which is where F0.5 punishes any guess.

A few more rules shaped the design:

And one rule of my own. The competition ran for three days, but I gave myself 24 hours, start to finish. That changes what “a good idea” means: anything that needs a second day of GPU time is out, every stage has to survive a crash without starting over, and there is no time to guess. Every choice below was measured on a holdout before it went in.

What the Data Looks Like

The data is synthetic: a real-looking business with noise applied to it several times. Here are two true match groups from the training data:

S1  Payne Enterprises       3315 Fremont Street, Peoria, IL
S2  Payne Énterprises       3315 FREMONT ST, PEORIA, IL
S2  PAYNE-ENRTPRMISES       3315 FREMONT SAINT, PEORIA, IL
S3  Payne Etrepndiels       3315 Fremont St, Peoria, Illinois
S3  Payne Enterprises  LLC  Fremont St, Peoria, Illinois

S1  Raj Investments LLP        6(29), C.I.T. Colony, 2Nd Main Road Mylapore, Chennai, Tamil Nadu
S2  ராஜ் இன்வெஸ்ட்மெண்ட்ஸ் எல்எல்பி  6(29), C.I.T. COLONY, 2ND MAIN ROAD MYLAPORE, CHENNAI, Tamil Nadu
S3  Raj Investments எல்எல்பி     6(29), C.i.t. Colony, 2Nd Main Road Mylapore, Chennai, TN
S3  ராஜ் இன்வெஸ்ட்மெண்ட்ஸ் எல்எல்பி  6(29), C.i.t. Colony, 2Nd Main Road Mylapore, Chennai, தமிழ்நாடு

Typos, injected accents, “St” expanded to “Saint”, a dropped house number, a legal form added, and for India, names and state names written in Tamil, Devanagari, Kannada and other scripts (28% of Indian Source 2 names). On top of that: junk around names (-- , << , | www.x.com), legal forms moved to the front (“LLC Moncada Léarning Center”), names written as a domain (maurewilliamscolombier.com), reordered address parts, null fillers, empty addresses, and records with a completely different trade name at the same address.

The hard part is not the noise, though. It is the distractors. About 26% of Source 2 and Source 3 records match nothing, and they are not random. They are near copies of a real business: a house number a few units off (1024 vs 1027), one real word swapped (First vs Seventh, Private vs Public), or a word like “& Sons” or “Group” added. In training, a record that adds “& Sons”, “& Associates”, “Group” or “Enterprises” to the business name at the same address is a distractor 98 to 100% of the time, while adding “Services” or “Center” is ordinary noise. And about half of all business names are shared by at least two different businesses in the same country, so a name alone is often not enough.

The Most Useful Fact

The training ground truth has 7,638,365 matched record IDs, and every single one is unique. Each Source 2 or Source 3 record belongs to at most one business.

That flips the problem around. Instead of only asking “which records match this business?”, you can ask “which business owns this record, if any?”. The pipeline uses this in three places: search runs in both directions (business to records and record to businesses), the models get features about how strongly other businesses compete for a record, and at the end each record is given to one business only.

Validation

Everything was measured on a split of the training businesses, and the split matters more than usual because the pipeline stacks models on top of models:

SplitShareUsed for
A80%training the bi-encoder and cross-encoders
B10%training the pruning model and the stacker
C18%calibration and choosing the final sets
C22%untouched, only for checking

The stacker has to learn from neural scores on businesses the neural models never saw, otherwise it learns to trust scores that are too good. Retrieval for B and C searches the full training pool, distractors included, so the holdout sees the same kind of search the test set does.

France has no labels, so I used a stand-in: train on US only and score India, and the other way round. Choices that helped US and India but hurt this cross-country transfer were left out.

The Pipeline

Pipeline stages from raw records to final matches, with candidate counts per business at each stage 1. Normalize transliterate, clean, mined alias maps ~12M records per split, all sources 2. Bi-encoder kNN 4 views, both directions, per country 58 candidates per business, 99.92% recall 3. Pruning model LightGBM on cheap scores 5.8 candidates per business, 99.82% recall 4. Pair scoring string features + 2 cross-encoders 5. Two-stage stacker LightGBM, stage 2 adds group features 6. Final sets one-to-one + expected F0.5 ~3.4 matches per business, precision 0.9993
The six stages, with the numbers measured on the holdout (split C1). The output of stage 3 is the candidate file, and it is exactly what stages 4 to 6 score.

1. Normalization

Every record gets a cleaned version next to the raw one: Unicode NFKC, lowercase, accents stripped, non-Latin scripts transliterated with anyascii, junk prefixes and suffixes removed, null-style fillers dropped, and repeated words collapsed. Legal forms (LLC, Inc, Pvt Ltd, SARL, SAS and so on) are found anywhere in the name and moved to their own field, which leaves a “core name”. Aliases (“X dba Y”, “t/a”, “formerly”) and domain stems get their own fields too.

Transliteration alone is not enough. लिमिटेड comes out as something like limitedd and Tamil Nadu comes out as tmilnatu. So instead of writing a dictionary by hand, the pipeline mines one from the training pairs: it lines up tokens between records that are known to match and keeps the mappings that show up consistently. That gives a 555-entry native-to-Latin word map, state aliases (tmilnatu and tn to tamil nadu) and street fixes (aveune to avenue). Some real entries from the mined maps:

After transliterationMined asKind
ailailpillpnative-script word
aimtrpraij'ij'enterprisesnative-script word
aiksportsexportsnative-script word
ainrjienergynative-script word
bombay, calcuttamumbai, kolkatacity alias
flor, fltfloor, flataddress typo
aveune, ausinavenue, austinaddress typo

All mined maps are keyed by the country string, so nothing is hard-coded to US or India. A short hand-written list covers French street abbreviations and maps French business words to English ones (“& Fils” to “& Sons”, “Groupe” to “Group”), so the distractor patterns the models learned in English also apply to French names.

2. Candidate Generation

Retrieval uses intfloat/multilingual-e5-small fine-tuned as a bi-encoder on split A. Round one uses in-batch negatives with batches of 1,024 pairs. Each batch holds one country and unique businesses, so no in-batch negative is ever a true match. Round two continues from round one with hard negatives mined from its own search results, and trains three views of each record: the full record (60% of batches), the name only (20%) and the address only (20%). The views matter because some records have no address, and some have a totally different trade name at the same address.

Every record is encoded in four views (round-two full, name and address, plus round one), and exact kNN runs on the GPU inside each country. It runs in both directions: each business looks for its top records, and each record looks for its top businesses. Because each record has one owner, the reverse direction is a strong filter. Two exact keys catch what embeddings miss: an acronym of the core name (for records like GAS) and a domain stem equal to the core name without spaces.

The union keeps 99.92% of true pairs but at 58 candidates per business, which is far too many. A small LightGBM model trained on split B cuts it down using only cheap signals: the cosine and both ranks of every view, four fast rapidfuzz scores, and rank and gap context. Pairs scoring at least 0.002 are kept, at most 15 per business:

True pairs kept by blocking against candidates per business, for each pruning threshold 97% 98% 99% 100% 3 4 5 6 7 used: 5.8 candidates, 99.82% 3.4 candidates, 97.4% candidates per business
Each dot is one pruning threshold, from 0.5 on the left to 0.0001 on the right. Past about 6 candidates, each extra candidate buys almost nothing. The union before pruning sits far off this chart: 99.92% at 58 candidates.

The output of this model is the candidate file, and it is exactly the set the rest of the pipeline scores. For comparison, the same e5 model without fine-tuning needs 60 candidates per business to reach 98.7%.

3. Scoring Pairs

Each candidate pair gets a wide set of features:

On top of that there are two cross-encoders, both microsoft/mdeberta-v3-base (278M parameters). One reads the cleaned text and one reads the raw text, since the raw version still has signals that cleaning throws away. Each is trained for one epoch on 3 million split-A candidate pairs, 40% of them positive, with the negatives coming from the retrieval lists so they are hard. One practical note: recent transformers versions load mdeberta-v3 in fp16, and training it that way gives NaN losses. Loading it in fp32 and training with bf16 autocast works.

So what does the stage-1 stacker actually lean on? Grouping its 131 features by where they come from:

Share of the stage-1 stacker's total gain by feature family Pruning score 67.8% Cross-encoder, raw text 16.5% Cross-encoder, cleaned 10.6% Cross-encoder context 2.5% Other 114 features 2.7%
Share of the stage-1 LightGBM model’s total split gain, clean run (the second run gives the same split to within a point). The pruning score already combines the bi-encoder cosines with four fast string scores, so the other 114 hand-built features each add little on their own.

4. Stacking With Group Features

A LightGBM stacker, trained 4-fold on split B, combines everything. Then a second stage adds group features, which come from one observation: true records agree with each other, and a distractor has a change nobody else shares. If four records of a business all say house number 1024 and a fifth says 1027, the fifth one is suspicious, even if 1027 is close to what the business itself says.

So for every candidate, stage 2 looks at the business’s other confident records from stage 1 (out-of-fold predictions, so there is no leakage) and measures how similar the candidate is to them:

strong = p.filter((pl.col("p") >= 0.5) & pl.col("is_arg")).select("i1", pl.col("i2").alias("j"))
pairs = p.select("i1", "i2").join(strong, on="i1").filter(pl.col("i2") != pl.col("j"))
...
g = pairs.group_by("i1", "i2").agg(
    pl.col("rc").max().alias("g_rc_max"), pl.col("rc").mean().alias("g_rc_mean"),
    pl.col("rn").max().alias("g_rn_max"), pl.col("rn").mean().alias("g_rn_mean"),
    pl.col("ra").max().alias("g_ra_max"), pl.col("ra").mean().alias("g_ra_mean"),
    pl.col("rnum").filter(pl.col("rnum_ok")).mean().alias("g_num_support"),
    ...
)

rc, rn and ra are the embedding, name and address similarity between the candidate and each confident sibling, and g_num_support is the share of siblings with the same house number. There is also a feature for whether the business’s own house number agrees with its records, which tells the model when the reference record is the noisy one. On the record side, stage 2 adds the probability of the best other business for this record and the margin to it.

5. Picking the Final Sets

The stacker gives a probability for each pair. Turning those into lists takes three steps. First, one-to-one: each record is kept only for the business where it scores highest.

def one_to_one(df):
    return df.sort(["i2", "p", "i1"], descending=[False, True, False]).unique("i2", keep="first", maintain_order=True)

Second, the probabilities are calibrated with isotonic regression on C1. Third, instead of one global threshold, each business gets the set that maximizes its expected F0.5. Sort its candidates by probability q. For the top-k set, the expected number of true matches in the set is roughly the sum of their q, and the expected number of true matches overall is the sum of all q. Plug those into the F0.5 formula and try every k. The empty set is also a choice: it scores 1.0 only if there is no true match at all, which has probability ∏(1 − q).

d = df.sort(["i1", "q"], descending=[False, True]).with_columns(
    pl.col("q").cum_sum().over("i1").alias("cs"),
    pl.col("q").cum_count().over("i1").alias("k"),
    pl.col("q").sum().over("i1").alias("tot"),
    (1 - pl.col("q")).log().sum().over("i1").exp().alias("p_empty"),
)
d = d.with_columns((1.25 * pl.col("cs") / (pl.col("k") + 0.25 * pl.col("tot"))).alias("ef"))

This handles businesses with no match without any special case: when every candidate is weak, the empty set wins. There is also an exact version that computes the full Poisson-binomial expectation on the GPU. Which one to use is decided on C1, and they differ by less than 0.0001.

6. Two Runs, and France

The whole pipeline (bi-encoders, cross-encoders, pruning model, stackers) is trained twice with different seeds. For US and India, the stage-2 logits of the two runs are averaged. A pair that one run did not keep as a candidate counts as probability 0.0001 in that run, and the candidate file is the union of both runs’ candidates, since that is the set the average scores.

France is handled differently. Two runs trained the same way disagree on 8% of French output rows, against 0.8% for US and India. The pairs only one run accepts sit right around the point where adding a pair stops paying off under F0.5, and there are no labels to check them against. So for France the final file keeps only the pairs both runs accept:

pp = p1.join(p2, on=["source1_entity_id", "matched_entity_ids"]) if a.mode == "intersect" else pl.concat([p1, p2]).unique()

Results

Macro F0.5
Final overall score (all test countries)0.988216
Holdout C1 (US and India)0.99254
Holdout C2, untouched (US and India)0.99219
US, C1 / C20.9917 / 0.9914
India, C1 / C20.9937 / 0.9934

On the holdout, pair-level precision is 0.9993 and recall is 0.977. The pipeline is very careful: it would rather miss a record than claim a wrong one, which is what F0.5 asks for. On the test set, blocking produces about 6.5 candidates per business per run.

The two cross-country checks (train on one country, score the other) come out at 0.986 for US to India and 0.982 for India to US. The fine-tuned bi-encoder features transfer especially well.

Where It Still Fails

Legal-form distractors. The most common wrong merge is a distractor that is the business name plus “Inc” or “LLC” at the same address. True records get an added legal form just as often, and sibling records share it about equally often in both cases (4.2% vs 4.0%). I could not find a signal that separates them.

Records with no address. True pairs with an address are recovered 99.8% of the time. Pairs where the record has no address, only 64%. Most of the misses have a name shared by several businesses in the same country, and the best weak signal I found (the right business has fewer records from the same source) is right only 39% of the time. With F0.5, predicting nothing is the correct call there.

France. The holdout says about 0.992 for US and India, and the accepted US and India test pairs look just like the holdout ones (same rates of word swaps, number changes, legal-form changes and empty addresses, same mean probability, same 3.4 matches per business). That puts France at roughly 0.96, and it is where most of the leaderboard gap comes from. French addresses end with a long region name that records drop 65% of the time or replace with a department (“Nord”, “Gironde”), and 13% of French businesses share an exact address with another one, against 4 to 6% in training. The French errors are mostly pairs the model is confidently wrong about, not threshold choices.

If I had more time, these are the things I would try next:

Where the 24 Hours Go

One full run of the pipeline, from raw TSVs to the two output files, takes just under 8 hours on one GPU. The final output averages two of them, so that is most of a day of GPU time on its own:

Minutes per stage for one full run on one GPU Load + normalize 3m Bi-encoder round 1 1h 09m Bi-encoder round 2 2h 00m Cross-encoder training 1h 20m Union + pruning 46m Pair features 9m Cross-encoder inference 2h 10m Stackers + groups 12m
Minutes per stage for one clean run, taken from its log. The blue bars are transformer training and inference, 85% of the 7h 49m total. Everything else, including all of the LightGBM work, fits in about 70 minutes.

That is why every stage writes its output to disk and skips work that is already done. If a job dies halfway, rerunning the same command picks up from the last finished stage instead of costing hours.

Try It

You will need the competition dataset, a Linux machine with one GPU with at least 40 GB of memory, about 100 GB of RAM and 150 GB of free disk.

git clone https://github.com/ArneshBanerjee/amazon-ml-challenge-2026-entity-resolution
cd amazon-ml-challenge-2026-entity-resolution
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements.txt \
    --extra-index-url https://download.pytorch.org/whl/cu126 --index-strategy unsafe-best-match

# two runs, averaged (the final output)
PY=$PWD/.venv/bin/python bash src/run_ensemble.sh /path/to/dataset /path/to/output

One run takes 8 to 10 hours on one GPU, most of it transformer training and inference. Every stage writes its output to disk and skips work that is already done, so if the machine goes down you can run the same command again and it continues where it stopped. run_all.sh does a single run in about half the time.

Credits

The challenge and dataset are from the Amazon ML Challenge 2026. The models are intfloat/multilingual-e5-small and microsoft/mdeberta-v3-base, both MIT licensed, along with LightGBM, polars, rapidfuzz, PyTorch and transformers. No external data was used.

Code: github.com/ArneshBanerjee/amazon-ml-challenge-2026-entity-resolution (MIT)
Author: arneshbanerjee.dev · Contact:

Stats

Loading views…

Sign in with GitHub, Google, or Discord to leave a comment.

Loading comments…

← back to homepage