Summary: A good teacher goes a long way! I distilled the cross-encoder from Part 3 into the Snowflake Arctic Embed L v2.0 bi-encoder (see Part 2). I used one contrastive warm-up round on Claude’s labels followed by three rounds of mining and scoring candidates with the teacher for training with a KL objective. The student steadily converged on the teacher with Spearman $\rho$ rising from 0.49 to 0.66, about 83% of the teacher’s 0.80. Retrieval metrics like ruling-level recall@10 improved less so (+0.055) and there appeared to be a divergence between rank correlation and recall. However, these results need to be considered exploratory since no cross validation was performed.
The series:
- Part 1: the Swiss legal system and the retrieval pipeline.
- Part 2: a heuristic benchmark built from commentary citations, comparing pretrained embedding models.
- Part 3: LLM relevance labels and a cross-encoder reranker fine-tuned with ordinal regression.
- Part 4 (this post): distilling the reranker into the first-stage retriever.
The Problem: Fast Retrieval vs. Accurate Ranking
Often there is a trade-off between doing things fast and doing them thoroughly. As discussed in Part 3, this also applies to retrieval pipelines. Bi-encoders enable fast scans of the database but cannot match the context sensitivity of the full self attention in cross-encoders. However, cross-encoders are more computationally intensive and do not readily scale to large sets of data.

The compromise made in practice is to retrieve documents in two stages. First, the bi-encoder does a rough pre-selection, casting a wide net that ensures most relevant documents are retrieved, even if not ranked properly. In a second step, the cross-encoder takes care of ranking the subset properly.
As discussed in Part 3, the cross-encoder was fine-tuned first since it is more sample-efficient to train. The final model was legal-swiss-roberta-large trained with CORAL ordinal regression on around 41k relevance-labeled pairs of queries and passages from rulings of the Swiss Federal Supreme Court as judged by Claude Sonnet 4.6.
In this part, I outline how I used the fine-tuned cross-encoder as a teacher for the bi-encoder selected based on the heuristic benchmark (Snowflake Arctic Embed L v2.0) in Part 2. Following Hofstätter et al. (2020)1 and Thakur et al. (2021)2, the goal is to transfer the cross-encoder’s relevance judgment to the bi-encoder so the retrieval stage surfaces better candidates for reranking.
Overview
- Corpus: 6'682 Swiss Federal Court rulings, 122'398 paragraphs of considerations ("Erwägungen") (see Part 1)
- Evaluation: 1'466 queries from 185 legal commentary articles3, 5-level relevance labels (0–4) (see Part 3)
- Bi-encoder student:
Snowflake/snowflake-arctic-embed-l-v2.0(see Part 2) - Cross-encoder teacher:
legal-swiss-roberta-largefine-tuned on 41k Claude-labeled pairs ($\rho$ = 0.7958) (see Part 3)
Fine-Tuning
Training Data
The dataset for fine-tuning the bi-encoder was constructed iteratively over multiple rounds. Concretely, after each round of fine-tuning, for each query the top 50 to 100 candidates4 were retrieved with the bi-encoder. Any new passages were labeled in relevance using the cross-encoder as a teacher and added to the training set. In the final round, this produced a dataset consisting of ~41k pairs labeled by Claude and 152k pairs labeled by the teacher model.
Fine-tuning followed the same procedure as detailed in Part 3. In particular, an ~80/20 train-validation split was used and data was strictly split on the article to avoid leakage.
| Source | Pairs | Score type |
|---|---|---|
| Claude Sonnet 4.6 | 41,438 | Ordinal labels $0–4$ |
| Cross-encoder teacher | 152,035 | Continuous $\mathbb E[Y] \in [0, 4]$ |
| Total | 193,473 |
Iterative Training Loop
Round 1: Contrastive Learning
The first round of training was meant to act as a warm-up for domain adaptation since the student model had no prior exposure to Swiss legal German. If it retrieves badly, the teacher’s scores over those candidates are nearly uniform and the distillation loss carries little information.
I used InfoNCE5 as the loss and treated label $\geq 3$ as positive and $\leq 1$ as negative. Paragraphs with label $2$ were discarded so the model is not penalized for failing to push away partially relevant passages.
InfoNCE treats ranking as classification. Given a query, its positive passage and $m$ negatives, the loss is the negative log-likelihood of picking the positive: $-\log \frac{e^{\text{sim}(q,p^+)/\tau}}{e^{\text{sim}(q,p^+)/\tau} + \sum_j e^{\text{sim}(q,p^-_j)/\tau}}$. Since the student scores with cosine similarity, all scores sit in a narrow band. Adding the temperature $\tau$ term helps spread them far enough apart for the softmax to discriminate. I set $\tau = 0.05$ following the default value in the sentence-transformers library for the contrastive loss.6
Round 2–4: KL-Divergence
After the warm-up round, I switched to minimizing the KL divergence between the teacher and the student’s relevance judgments. No binarization is necessary here, since all paragraphs produce training signal. The two distributions are formed over the candidate list of each query. For a query $q$ with candidates $p_1, \dots, p_n$, the teacher supplies $t_i = \mathbb{E}[Y\mid q, p_i]$ and the student supplies the cosine similarity between query and candidate $s_i$. Both are standardized to obtain $\tilde t_i$ and $\tilde s_i$ before applying a softmax: $$P_i = \text{softmax}\left(\tilde{t}_i / \tau\right), \qquad Q_i = \text{softmax}\left(\tilde{s}_i / \tau\right)$$
The student is hence trained to produce embeddings as to spread the probability mass of relevance the same way the teacher does. What transfers is the relative spacing of the candidates in embedding space.
The Kullback–Leibler divergence measures how far one probability distribution sits from another. For two discrete distributions $P$ and $Q$ over the same $n$ outcomes: $$D_{\text{KL}}(P | Q) = \sum_{i=1}^{n} P_i \log \frac{P_i}{Q_i}$$ It is zero exactly when $P = Q$, and grows when $Q$ places little mass where $P$ places a lot. Note that it is not symmetric and that I use $P$ = teacher, $Q$ = student. The student is penalized hardest for ignoring a passage the teacher considers relevant and less so for hedging on one the teacher dismisses.
The training loop was structured as follows:
- Mining the top-k candidates from the full corpus for each training query using the current bi-encoder
- Scoring all newly retrieved candidates using the cross-encoder teacher, assigning each a continuous $\mathbb{E}[Y]$ score
- Training using the KL-divergence loss to align the student’s cosine similarity distribution with the teacher’s score distribution over all candidates per query
- Repeat with the updated bi-encoder
Results
Before discussing the numbers, I need to point out a caveat in the following results. No cross-validation was performed here but only a simple train-validation split with 37 held-out articles and 354 corresponding queries. Because queries cluster within commentary articles, sharing topic, vocabulary and cited rulings, the effective sample is nearer 37 than 354. The following are therefore exploratory findings. Due to a shift in the project’s focus, I did not follow them up with more detailed analyses. I explain more of my reasoning for the shift in focus in the conclusion below.
At a high level:
The student steadily converges on the teacher. Spearman $\rho$ rises from 0.4924 to 0.6589 across four rounds, roughly 80% of the teacher’s 0.7958, improving every round. The cross-encoder’s own view of what the student retrieves moves in accordance. Based on the cross encoder’s judgment, CE precision@5 rises from 0.148 to 0.211 and mean $\mathbb{E}[Y]$@5 from 1.756 to 2.067.
Retrieval improves but the evidence is thinner. Measured against the commentary citations, ruling-level recall@10 gains +0.055 from pretrained to Round 4 and bootstrapped confidence intervals show a positive effect. At the consideration level, the gain is +0.032. When using bootstrapped confidence intervals, the lower bound was barely positive and Label recall@10 (see definitions below), which gained +0.032, included zero in its interval.
Model Ruling@10 Erwägung@10 Erwägung@20 Label@10 Pretrained 0.385 0.323 0.380 0.637 Round 1 0.424 0.345 0.414 0.659 Round 2 0.433 0.358 0.422 0.672 Round 3 0.432 0.343 0.416 0.654 Round 4 0.439 0.355 0.411 0.669 How to read the columns: All four are recall at $k$, i.e., of the passages that should have been found for a query, what fraction actually appeared in the model’s top $k$ out of all 122,398 paragraphs?
- Ruling@k: Did any paragraph from the ruling cited by the commentary author enter the top $k$?
- Erwägung@k: Did the specific paragraph cited in the commentary enter the top $k$?
- Label@k: What fraction of the paragraphs Claude labeled $\geq 3$ for this query entered the top $k$?
Taken together, the later rounds kept improving rank correlation without improving recall, which was a surprising finding. Since Spearman is computed only on the ~30 labeled passages per query, it suggests the student learned to order those passages better but without generalizing to the rest of the corpus. The training setup may have encouraged this, since the teacher only ever scores candidates the student has already retrieved. However, given the lack of cross-validation and the size of the held-out set, a sampling effect cannot be ruled out. Even so, in hindsight, checkpoints should have been selected on held-out recall rather than rank correlation, since recall is what a first-stage retriever needs to optimize.
Conclusion
Despite the mentioned limitations, I consider these promising results. When I tested the models with my own queries on a bare bones frontend, I found the retrieval quality to be impressive. Relevant passages are retrieved and the assigned relevance scores appeared reasonable to me.
However, as soon as one queries for topics from areas of law not covered by the commentary, the model breaks and does not retrieve relevant rulings. On a positive note, the assigned relevance scores for the retrieved rulings are also very low. The interpretability of the cross-encoder’s predicted probabilities proved very useful in this regard, enabling an assessment of uncertainty.
Even so, this points to a larger issue with the search problem, namely access to training data. Unfortunately, I don’t have access to the extensive, high-quality legal commentaries, which would vastly improve the identified issues but require commercial licenses. Hence, I have decided to focus on a different area of legal tech, where training data is less of a concern.
I hope to report more on my new direction soon!
Appendix
The following numbers were all obtained on a held-out set of 354 queries
Spearman rank correlation
The contrastive warm-up is the largest single jump, and by per-query $\rho$ it beats the three distillation rounds combined. Round 4 adds little.
| Round | Loss | Train data | $\rho$ overall | $\rho$ per query | $\Delta\rho$ |
|---|---|---|---|---|---|
| Pretrained | - | - | 0.4924 | 0.3531 | - |
| 1 | InfoNCE | 31.5k labeled | 0.5698 | 0.4502 | +0.0774 |
| 2 | KL | + 38.9k mined | 0.6142 | 0.4904 | +0.0444 |
| 3 | KL | + 97.2k mined | 0.6483 | 0.5113 | +0.0341 |
| 4 | KL | + 115.0k mined | 0.6589 | 0.5213 | +0.0106 |
The following shows the number of new candidates added per round through the mining process as described above. Note the sharp drop in round 3, presumably because the student increasingly re-retrieves passages that have already been scored.
| Mining run | $k$ | New pairs |
|---|---|---|
| 1 | 50 | 51,695 |
| 2 | 100 | 76,880 |
| 3 | 100 | 23,460 |
Retrieval Evaluation
Spearman measures ranking agreement on labeled pairs but does not necessarily imply good retrieval quality. What the pipeline depends on is whether the right passages appear in the top $k$ out of all 122,398.
I report two different metrics based on the source of the relevance judgment:
- Citation recall based on the specific rulings and Erwägungen cited by the commentary articles. Those citations were written by legal commentators with no LLM involved.
- Label recall, referring to recall for passages Claude labeled $\geq 3$.
Citation Recall@k (Erwägung level)
| Model | @5 | @10 | @20 | @50 | @100 |
|---|---|---|---|---|---|
| Pretrained | 0.2456 | 0.3232 | 0.3803 | 0.4319 | 0.4804 |
| Round 1 | 0.2631 | 0.3454 | 0.4138 | 0.4742 | 0.4913 |
| Round 2 | 0.2760 | 0.3584 | 0.4220 | 0.4826 | 0.5160 |
| Round 3 | 0.2766 | 0.3425 | 0.4160 | 0.4896 | 0.5141 |
| Round 4 | 0.2838 | 0.3550 | 0.4110 | 0.4909 | 0.5177 |
Label Recall@k (Claude label $\geq 3$)
| Model | @5 | @10 | @20 | @50 | @100 |
|---|---|---|---|---|---|
| Pretrained | 0.4620 | 0.6369 | 0.7890 | 0.9106 | 0.9428 |
| Round 1 | 0.5007 | 0.6586 | 0.7951 | 0.9078 | 0.9521 |
| Round 2 | 0.5051 | 0.6719 | 0.8002 | 0.9188 | 0.9589 |
| Round 3 | 0.4978 | 0.6544 | 0.8122 | 0.9164 | 0.9574 |
| Round 4 | 0.4956 | 0.6688 | 0.8072 | 0.9195 | 0.9612 |
Cross-Encoder Verification
Only pairs from the original training set carry labels from Claude. As the distilled model improves, it retrieves new paragraphs, which is also seen in the coverage of Claude labeled paragraphs dropping from 73% to 55% at $k=5$. To measure whether this represents generalization or drift, I scored every retrieved passage with the cross-encoder teacher. Under the teacher’s judgment, retrieval quality improves in every round and at every depth.
CE Precision@k (E[Y] ≥ 3, all passages scored)
| Model | @5 | @10 | @20 |
|---|---|---|---|
| Pretrained | 0.1480 | 0.1020 | 0.0658 |
| Round 1 | 0.1927 | 0.1339 | 0.0874 |
| Round 2 | 0.1983 | 0.1415 | 0.0908 |
| Round 3 | 0.2062 | 0.1492 | 0.0963 |
| Round 4 | 0.2107 | 0.1500 | 0.0987 |
CE Mean E[Y] (all passages scored)
| Model | @5 | @10 | @20 |
|---|---|---|---|
| Pretrained | 1.756 | 1.545 | 1.330 |
| Round 1 | 1.924 | 1.685 | 1.442 |
| Round 2 | 2.011 | 1.802 | 1.558 |
| Round 3 | 2.059 | 1.846 | 1.609 |
| Round 4 | 2.067 | 1.867 | 1.630 |
Labeled vs. Unlabeled: Is the Model Generalizing?
By splitting CE scores between paragraphs that have Claude generated labels and those that don’t, we can see whether the newly added retrievals are relevant. Indeed, the unlabeled retrievals improve from 1.72 to 1.99 at @5. This suggests that the student reproduces the teacher’s preferences on articles neither model trained on, rather than drifting toward passages the teacher considers irrelevant.
| Model | E[Y] labeled @5 | E[Y] unlabeled @5 | Unlabeled fraction @5 |
|---|---|---|---|
| Pretrained | 1.773 | 1.718 | 27.1% |
| Round 1 | 1.999 | 1.848 | 37.6% |
| Round 2 | 2.066 | 1.943 | 40.5% |
| Round 3 | 2.131 | 1.995 | 44.0% |
| Round 4 | 2.157 | 1.987 | 44.8% |
Hofstätter, Althammer, Schröder, Sertkan, Hanbury, 2020: “Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation”, https://arxiv.org/abs/2010.02666 ↩︎
Thakur, Reimers, Daxenberger, Gurevych, 2021: “Augmented SBERT: Data Augmentation Method for Improving Bi-Encoders for Pairwise Sentence Scoring Tasks”, https://arxiv.org/abs/2010.08240 ↩︎
Commentary articles from https://onlinekommentar.ch, see Part 1 for more details. ↩︎
k=50 during warm-up, k=100 thereafter. See Appendix for details. ↩︎
Karpukhin, Oguz, Min, Lewis, Wu, Edunov, Chen, Yih, 2020: “Dense Passage Retrieval for Open-Domain Question Answering”, https://arxiv.org/abs/2004.04906 ↩︎
Gao, Yao, Chen, 2021: “SimCSE: Simple Contrastive Learning of Sentence Embeddings”, https://arxiv.org/abs/2104.08821 ↩︎