Summary: I generated ~1'400 expert-style queries from legal commentary passages and used the rulings those passages cite as proxy relevance labels. Among three pretrained multilingual embedding models, Snowflake’s arctic-embed-l-v2.0 retrieves best, notably so on cross-lingual queries. Still, the cited ruling appears in the top 10 for fewer than half of the queries. Expanding results along the citation graph does not appear to help, but simply retrieving more candidates does, which motivates the reranking in Part 3.

The series:

  • Part 1: the Swiss legal system and the retrieval pipeline.
  • Part 2 (this post): a heuristic benchmark built from commentary citations, comparing pretrained embedding models.
  • Part 3: LLM relevance labels and a cross-encoder reranker fine-tuned with ordinal regression.
  • Part 4: distilling the reranker into the first-stage retriever.

Recap

Part 1 served as a crash course on the Swiss legal system. It discussed:

  • why Federal Supreme Court rulings (BGEs) are an important basis for legal research despite Switzerland not being a strict case-law jurisdiction.
  • the challenge posed by the lack of a clean ground truth measure for the relevance of a court ruling to a query.

Creating a Query-Set

Any retrieval starts with a query and so does building a retrieval system because what good is a system that cannot act on an actual user’s query (not much…).

To mimic real user queries and build the evaluation set, all commentaries from the Online-Kommentar were scraped and parsed into paragraphs.1 For each paragraph that cited a landmark (BGE identifier) ruling of the Federal Supreme Court, I used Claude Sonnet 4.6 to generate three queries with the following profiles:

  1. Expert: Mimicking how a lawyer or student would look up a legal question, mentioning the legal provision and legal terminology.
  2. Layman: Representing a layman’s query, excluding any article provisions or legal terminology.
  3. Keyword: Querying by keyword search to compare sparse retrieval with dense queries.

Claude was instructed:

  • To generate a query that a user would type to retrieve the given passage.
  • Not to reference the court ruling, commentary or author to prevent retrieval based on meta-data.
  • Not to paraphrase the passage but distill the core legal question being discussed.

In total, around 1'400 query-target pairs of each profile were generated and formed the basis for the evaluation benchmark.

Sample Queries

In preliminary testing, I found that the expert queries produced better success rates (around +5-10%) compared to the other two. The retrieval trends discussed below were otherwise very similar. Given that the target users are legal professionals and due to limited compute, I focused on the expert queries and left the other categories for a later point in time. All queries are in German.

To give you a sense of the expert query, here are some examples:

“Zuständigkeit zur Statutenänderung bei Familienstiftungen gemäss Art. 87 ZGB: Stiftungsrat oder Zivilgericht?”

“Umfang der Auskunftspflicht des Drittschuldners gegenüber dem Konkursamt gemäss Art. 222 SchKG in Analogie zu Art. 400 OR”

“Abgrenzung Meinungsfreiheit Art. 16 BV und Wirtschaftsfreiheit Art. 27 BV bei kommerziellen Äusserungen und Werbung”

Evaluation

Heuristics for Retrieval Quality

The question that immediately follows the creation of the queries is how to then evaluate the quality of retrieval. Given that I couldn’t manually go through all retrievals, I initially created a heuristic-based evaluation to get a sense of the retrieval quality that pre-trained embedding models would achieve while scaling to the entire set of queries.

Recall that for each paragraph, the cited court ruling was parsed. Now given that the ruling was used as a reference, it likely contains substantial evidence that supports the statement of the paragraph. If the commentary author used BGE 141 IV 108 to back their legal claim, a system retrieving BGE 141 IV 108 for a query about that claim is likely doing its job well. Therefore, I used the extracted citations as proxy relevance judgments.

Judging Relevance through Proxy-Metrics

Indeed, matches at the consideration level, i.e., when the system retrieves the exact ruling id and consideration number (see Part 1 for a crash-course on Swiss court rulings) proved to be a reliable signal both on a small, manually inspected sample and later LLM-based relevance judgments (discussed in Part 3).

However, multiple rulings may support a given statement. Even expert human-judgment would hence not be expected to yield exact matches in all cases for top-1 retrieval (though likely within the top 5-10). Unfortunately, each paragraph generally only cites a single ruling rather than a list of all acceptable ones, so this proxy alone only produces a very sparse signal.

Moreover, it cannot assess the topicality of a retrieved document. If the model retrieves a ruling that is closely related in content but cannot be used to support the statement, based on this proxy alone, one would equate it with a model that extracted no signal at all.

Since I wanted to measure the signal that is understood by the pre-trained general-purpose models in more detail, I developed additional proxies, which are listed below. For instance, the presence of article references should also serve as a proxy for the topicality of the retrieved documents.

Classification hierarchy (best to worst):

ClassDescriptionExample
exactSame ruling ID, and the Erwägung matches perfectly (direct equality, substring match, or falls within a cited range).Cited: E. 3.2
Retrieved: E. 3.2 (or 3.2.1)
siblingSame ruling ID, and shares the same top-level Erwägung numeral, but is not an exact match.Cited: E. 3.2
Retrieved: E. 3.5
partialSame ruling ID, but a completely different top-level Erwägung entirely.Cited: BGE 141 IV 108 E. 5
Retrieved: BGE 141 IV 108 E. 2
article_exactDifferent ruling, but the retrieved Erwägung directly cites the exact commentary article being queried.Query: “Art. 210 OR”
Retrieved: “Art. 210 OR”
article_baseDifferent ruling. Matches the base article number within the retrieved Erwägung, but differs in its Latin suffix.Query: “Art. 322quinquies StGB”
Retrieved: “Art. 322ter StGB”
article_rulingDifferent ruling. The retrieved Erwägung doesn’t cite the article, but the ruling as a whole does (substring match).Query: “Art. 52 ZGB”
Retrieved paragraph: no reference to Art. 52 ZGB
Elsewhere in the ruling: “Art. 52 Abs. 1 ZGB”
law_matchDifferent ruling. The retrieved text cites the same law abbreviation, but a completely different article number.Query: “Art. 210 OR”
Retrieved: “Art. 97 OR”
shared_articlesDifferent ruling. The specifically cited ruling and the retrieved ruling share >= 2 substantive (non-ubiquitous) articles.Both rulings cite:
“Art. 41 OR” & “Art. 42 OR”
missNone of the above matching criteria are met.Completely unrelated ruling and articles.

Evaluation Pipeline

This gives us the following pipeline as the basis for the evaluations in this first exploratory phase:


Pre-Computed: 
                Commentary passage → Query (Claude API)

---

At test-time: 
                Expert-Query -> ChromaDB (cosine distance) 
                                            ↓             
                                    Top-k Erwägungen
                                            ↓
                                Classification & Evaluation

Limitations

As discussed below, it is important to note that a significant portion of the misses were found to be on-topic during manual inspection but still fell through the above classification. A subsequent analysis of the heuristic benchmark with an LLM relevance judgment confirmed this and pointed out gaps in the heuristic benchmark.

Since these gaps apply equally to all models, it’s reasonable to assume that they do not substantially affect the relative ranking of the models, so the findings of the heuristic benchmark should still be instructive.

Results

Specifications

I first benchmarked the following three embedding models, all of which have multi-language support and are openly available via the HuggingFace sentence-transformer library:

name#paramsdimscontextdescription
intfloat/multilingual-e5-large~560M1024512Robust baseline for semantic search tasks by Microsoft, trained on over 100 languages.
mixedbread-ai/deepset-mxbai-embed-de-large-v1~560M1024512Collaboration between deepset and Mixedbread AI built on multilingual-e5-large but fine-tuned with 30 million German data pairs.
Snowflake/snowflake-arctic-embed-l-v2.0568M10248192Built on BGE M3 and finetuned by Snowflake for multilanguage performance and with a large context window.

Analyzed Aspects

The following discusses the retrieval quality by aspect that was analyzed. All of these results are based on the heuristic classification benchmark described above and compare the retrieval quality of the pre-trained embedding models. Below I go over some metrics that should be indicative, including retrieval success rates and some metrics on the embedding space. No additional retrieval signal like BM25 or cross-encoder reranking was used in the retrieval for these results.

Success@K

Success@k measures the share of queries that retrieve a document of at least the indicated relevance or better among the top-k results. A success@10 of 50% for ‘Any BGE’ means that for half of all queries, the cited ruling, though not necessarily the correct consideration number, was among the 10 results returned.

The plots group the match classes from the table above into cumulative levels, where each level includes all better ones:

  • Exact: exact only, i.e., the cited consideration was retrieved.
  • Any BGE: exact, sibling or partial, i.e., any paragraph of the cited ruling was retrieved.
  • Strong article: any of the above, or article_exact, i.e., the retrieved paragraph cites the queried article.

At a high level, the plots below show that success is modest, staying below 50% for all three models at the ruling level (Any BGE), with arctic reaching 47.5% at k=10. In other words, for more than half of all queries, no paragraph from the cited ruling makes it into the top 10.

Two further trends are visible:

  1. Success rate for strong article is quite high for all models, indicating that topical rulings are recovered. This highlights the heuristic’s blind spot for semantic relevance, a limitation that is later confirmed by the LLM-based judgments.
  2. Arctic outperforms e5-large and the deepset model across all levels.

Success@k

Language Comparison

As a follow-up, I compared the success rate between French (~30% of corpus) and German target rulings (~66% of corpus), i.e., distinguishing by the language of the ruling cited in the commentary. Since all queries are German, the French split is a cross-lingual retrieval task.

All models perform significantly worse in this cross-lingual set-up. Notably, e5-large and deepset collapse almost entirely, hardly ever retrieving the targeted French ruling for any of the corresponding queries.

Success@k

For 157 commentary paragraphs, both a German and a French ruling were cited. Comparing retrieval on these samples holds the query fixed. The two cited rulings may still discuss different aspects though, so this is only an exploratory finding on the cross-lingual effect. The table counts, at k=10, for how many of these 157 queries each model retrieved both rulings, only the German one, only the French one, or neither:

namebothDE onlyFR onlyneither
e5_large097060
deepset2102053
arctic1687351

The results suggest that arctic is clearly better at cross-lingual retrieval than the other two, which is important in the legal context. Even so, arctic also suffered a drastic loss in performance such that cross-lingual retrieval should be analyzed more extensively in future work. In particular, it would be interesting to see whether translating queries and rulings into a common language closes the gap.

Analyzing the Embedding Space

In the tested baseline retrieval system, the retrieval quality depends entirely on how well the embedding model picks up on the topical differences in the documents and queries. The following results serve to understand the embedding space in more detail.

Ideally, the cosine distance in embedding space would reflect topical divide and map different topics far apart in the space of embeddings, resulting in a large cosine distance (a large angle between the vectors). However, all three of the models are general-purpose models that have been trained on many languages, of which the Swiss legal domain is only a subset of a subset. Therefore, clear separation cannot be expected.

Distribution of Cosine Distance by Match

Indeed, the results below show a large overlap in the distribution of the cosine distances between the query and retrieved document for different classes. This suggests that the models have a hard time separating relevant from non-relevant documents.

distance_distribution

Comparing only rulings hits (exact, sibling, partial) vs. misses on the German target group, the overlap of the cosine distances is apparent visually.2

cosine_distribution_hit_vs_miss

Cohen’s d quantifies this by measuring the difference between the two means relative to their pooled standard deviation. On the German target group, all three models reach values of 0.7–0.85, generally considered moderate to large. Hence, on average, the models do separate hits from misses but the distributions still overlap significantly. arctic achieves the highest separation.

ModelHit meanMiss meanGapCohen’s d
e5_large0.12180.13210.01020.698
deepset0.15550.17090.01540.690
arctic0.36650.42130.05480.849

Note that the absolute values are not comparable, only the Cohen’s d values are. Hits = exact/sibling/partial. Misses = match_type ‘miss’ only.

However, the table only compares hits (match type exact/sibling/partial) and misses (match type miss), leaving out the middle ground of rulings that the heuristic classifies as topically related (e.g., based on sharing an article) but would nonetheless not be considered relevant in practice. These are the hard cases and likely sit much closer to the hits in embedding space, so the separation a retriever actually faces is probably weaker than suggested here.

At the same time, as mentioned before, manual inspection revealed that the heuristic also misses relevant rulings, which inflates the measured overlap. Two factors promote this:

  • First, the ground truth is incomplete since a commentary paragraph cites one ruling but multiple equally valid rulings exist on the same topic.
  • Second, the article matching in the classification pipeline cannot detect connections when no references are made and it cannot detect connections relying on semantics. For instance, two rulings on inheritance law citing Art. 527 vs. Art. 636 ZGB are clearly related but share no article references.

Both limitations demand caution in interpreting these results and motivate the LLM judgments used in Part 3 to measure discriminative power more precisely.

Distance to Correctly Retrieved Rulings, German vs. French targets

Lastly, the strong effect of language noted earlier in the success rates also manifests here. Comparing the cosine distance between the query and the target ruling for all retrieved hits (matching BGE), the arctic plot shows a strong shift of the French target group. Language thus appears to be a strong factor in the embedding here and may cause topically relevant French documents to rank below irrelevant German documents driven by language alone.

cosine_distribution_lang

Citation Graph Expansion vs. Deep Retrieval

As a final experiment, I looked at whether the citation network between rulings provides any additional retrieval signal that dense embeddings miss. Intuitively, if a retrieved consideration cites a past ruling (outgoing citation) or is cited by a newer ruling (incoming citation), those connected documents should be relevant given the limited scope of a paragraph.

To test this, I implemented a graph expansion which fetches the citation neighbors of the top-k results at query-time, scores all paragraphs of those neighbors using cosine distance and extracts only the Top-3 most relevant paragraphs per neighbor. This is to prevent the candidate pool from exploding with irrelevant text.

However, expanding the candidate pool inherently increases the chances of finding the target ruling. For a fair comparison, I evaluated pure dense retrieval at the same number of retrieved candidates in the expanded set per query (N). For instance, if graph expansion yielded 111 candidates, I compared it against dense Success@111. All numbers below are for arctic at the Any BGE level and Lift is the difference between graph expansion and Dense@N.

As the results below show, for both citation directions, graph expansion underperforms pure dense retrieval at matched candidate depth. Note that the penalty is notably larger for incoming citations, likely a consequence of the larger fanout. Widely cited rulings accumulate many incoming citations, so incoming expansion produces a deeper candidate pool (avg. N = 92.4 at k=10 vs. 32.2 for outgoing) and is therefore matched against a correspondingly deeper dense control.

At the same time, it’s noteworthy that the Dense@N results show a strong lift in retrieval success up to around N=100. Accordingly, reranking using a more context-sensitive model, like a cross-encoder, could help filter the relevant results better. Overall, these results suggest that citation-graph expansion offers no retrieval gain over simply searching deeper.

  • Expansion with outgoing citations

    kBaseGraph (Avg N)Dense@NLift
    120.9%28.5% ( 5.3)35.0%-6.5%
    331.3%42.3% ( 12.2)46.7%-4.5%
    538.1%49.6% ( 18.4)53.5%-3.9%
    1047.5%58.7% ( 32.2)60.2%-1.5%
  • Expansion with incoming citations

    kBaseGraph (Avg N)Dense@NLift
    120.9%27.5% ( 15.0)44.2%-16.7%
    331.3%39.7% ( 35.8)58.8%-19.1%
    538.1%46.9% ( 53.1)63.8%-16.9%
    1047.5%54.4% ( 92.4)70.9%-16.6%

Next Steps

The heuristic benchmark provides a solid baseline understanding of the pretrained embeddings and identifies Snowflake’s arctic model as the strongest of the three tested models. It also points to three areas for improvement:

  1. Moving from heuristics to LLM judgments. The heuristic systematically underestimates true success because it cannot detect semantic relevance without explicit article references. Following the research on LLMs as judges, scoring relevance on a multi-level scale should give a more accurate signal. This matters most for training, since fine-tuning needs a reliable and ideally fine-grained signal.

  2. Reranking with a cross-encoder. At small k, success rates are low, but they rise steeply as k grows, reaching ~70% among the top 100 candidates. The relevant rulings are often retrieved but just not ranked high enough, and the embedding-space analysis showed that bi-encoders struggle to separate hits from misses. The standard remedy is to retrieve a pool large enough for good recall and let a cross-encoder rerank it. Cross-encoders process query and document jointly, which allows a much finer-grained relevance judgment than comparing two independently computed embeddings. Part 3 covers this in detail.

  3. Investigating the Cross-Language Gap. The stark drop in success rates for French rulings when using German queries is a bottleneck given that 30% of the corpus is French. Future work should test whether translating query and document into the same language prior to retrieval closes this gap. In any case, an effective retrieval system will need to deal with this issue. Otherwise a large part of the corpus is effectively inaccessible.


  1. See Part 1 for an explanation of legal commentaries and their significance. ↩︎

  2. The French target group was omitted to avoid language as a confounding factor. ↩︎