Summary: I built a retrieve-and-rerank search system over ~6'600 landmark rulings of the Swiss Federal Supreme Court (122k paragraphs, German/French/Italian). Due to the lack of human relevance labels, I used citations in legal commentaries as a relevance proxy. Later, I used an LLM to create a synthetic dataset for finetuning, letting it judge relevance for ~41k query-document pairs. A cross-encoder pretrained on Swiss legal text and fine-tuned with ordinal regression (CORAL) reached a Spearman $\rho$ of ~0.80 against those judgments with calibrated inclusion probabilities.

The series:

  • Part 1 (this post): the Swiss legal system and the retrieval pipeline.
  • Part 2: a heuristic benchmark built from commentary citations, comparing pretrained embedding models.
  • Part 3: LLM relevance labels and a cross-encoder reranker fine-tuned with ordinal regression.
  • Part 4: distilling the reranker into the first-stage retriever.

Introduction

In this first blog-post, I broadly outline how the Swiss legal system works, what its main sources are and how these are structured. I will try to keep this light-hearted while also giving you a sense of the information needs in Swiss law.

In particular, I discuss why court rulings by the Federal Supreme Court should form a good basis for this first stretch of the project and what data one can hope to extract from a court ruling. Finally, I discuss the implications for retrieval systems and the hurdles for evaluation. The technical implementation of the evaluation and first benchmarks are discussed in the next post.

Whereas the United States and the United Kingdom have a case-law jurisdiction, Switzerland, like most countries in Continental Europe, has a codified legal system. At its core, this means that Swiss courts are not bound by precedents set in earlier rulings by other courts. Instead, every court applies the law by interpreting the provisions of the legal code, which form the basis of any legal discussion.

The following is a brief crash course of the Swiss legal system.

A Crash-Course

We can broadly divide the law into two areas:

  1. Public Law, which deals with the law governing the relation between the state and all private parties. Among others, this includes:
    • Administrative Law, governing the rights and duties of the state administration, since the state always needs a legal basis to act.
    • Criminal Law (so a parking ticket doesn’t send you to prison)
    • Tax Law, and so on.
  2. Private Law, which deals with the law governing the rights and duties of private parties. Among others, this includes:
    • The Civil Code (ZGB), which defines the very basic notions of legal personhood but also deals with matters of family law, inheritance law and property law.
    • The Code of Obligations (OR), which guarantees private parties the right to form contracts and outlines the legal framework for it, protecting private persons while also giving them considerable freedom.
    • Company Law (part of the OR), which provides the legal forms, such as the stock corporation or the limited liability company, through which persons can join together for a common purpose.

You get the idea. As cruel as it sounds, the goal is to govern every part of life in a codified manner. Specifically, every legal norm is required to be “general” and “abstract” in its nature, meaning it applies to (i) an indeterminate number of persons (not to any one or group of persons specifically) and (ii) applies to an indeterminate number of cases.

To illustrate this, consider Art. 2 ZGB (article 2 of the Civil Code):

Art. 2 para. 1 ZGB: Every person must act in good faith in the exercise of his or her rights and in the performance of his or her obligations.

I hope you agree that this does not lack in abstractness or generality. I also hope you see the problem that is involved here:

  • Who counts as a person? Does a company limited by stock count as a person?
  • What does “good faith” encompass?
  • What are “rights” and what are “obligations”?
  • Can I ever not act in good faith? (note, you shouldn’t)

Making matters worse, the principle of good faith that follows from Art. 2 ZGB is so fundamental, that many other provisions reference it. For instance, consider Art. 23 OR and Art. 25 OR (Code of Obligations), which state:

Art. 23 OR: A party labouring under a fundamental error when entering into a contract is not bound by that contract.

Art. 25 para. 1 OR: A person may not invoke error in a manner contrary to good faith.

So we’re relying on Art. 2 ZGB to restrict the set of reasons allowed for backing out of a contract.

Norm Interpretation

All this goes to say that we need a system to interpret the law well and consistently (this principle is also codified; who would’ve thought). Luckily, a stable set of ideas and practices has developed, which the courts abide by. Notably, when the court applies the law, it considers and combines four approaches:

  1. Grammatical approach, interpreting a norm according to the ordinary use of language, terms, grammar and syntax
  2. Systematic approach, interpreting a norm in its systematic context, i.e., the provisions and sections surrounding it in the code.
  3. Historical approach, which aims to interpret a provision in light of (i) the intent of its drafters and (ii) the meaning attributed to it by society at the time of its enactment.
  4. Teleological approach, which searches for the purpose of a norm and interprets them in the light of the purpose, values, legal, social and economic goals those provisions aim to achieve.

Where various interpretations are at hand, the one more in line with the constitution is preferable and conflicts with international law need to be considered.

Luckily, the legal process is well-documented. Every revision of federal law is accompanied by a publicly available comment on the reasoning and justification behind new law. Furthermore, practitioners and academics write legal commentaries, which are detailed discussions and summaries of the legal practice for each individual legal norm.

The Federal Supreme Court is the highest court in Switzerland and has the final say in the application of the law. Again, contrary to case-law jurisdictions, the Federal Supreme Court decides on a case-by-case basis and is in theory not bound by any precedents. However, in practice, it needs to ensure that it maintains consistency to guarantee equality under law and prevent arbitrariness. Therefore, decisions by the Federal Supreme Court still carry a very large weight in practice and form an important legal source for practitioners.

In this project, I will therefore focus on developing a retrieval pipeline for rulings of the Swiss Federal Supreme Court.

Court Rulings as a Datasource

Rulings of the Swiss Federal Supreme Court luckily come in a standardized format and it makes sense to briefly comment on the structure and content for readers that do not have a legal background.

Every published ruling by the Swiss Federal Supreme Court is assigned an ID, which provides useful information. For instance, in BGE 141 IV 108:

  • the BGE means that this is an officially published, landmark ruling
  • 141 is the volume and encodes the year of publication (141 + 1874 = 2015)
  • IV corresponds to the Criminal Law department. The departments are classified as:
    • I (Constitutional/Public Law),
    • II (Administrative Law),
    • III (Civil Law, Debt Collection, Bankruptcy),
    • IV (Criminal Law), and
    • V (Social Security Law).
  • 108 is the page number.

Structurally, every ruling consists of:

  • A keyword summary (“Regeste”) of the discussed legal aspects and legal provisions.
  • The facts of the case (“Sachverhalt”), where the relevant facts are summarized.
  • The considerations (“Erwägungen”), which comprise the court’s legal reasoning and judgment for each legal aspect in question. The considerations are grouped and numbered by the aspect in that section. They are generally self-contained in their treatment of the legal aspect they focus on.

Designing a Retrieval Pipeline

The objective of the retrieval pipeline is to retrieve the most relevant paragraphs from the considerations (“Erwägungen”) of the Federal Supreme Court for a given query. The corpus of court rulings spans published rulings in the official collection, an officially curated selection of the most important rulings, typically those answering a new question or changing existing practice. Published rulings are also condensed down to the considerations worth publishing. Hence, they should form a well-curated selection covering all areas of law. For practical reasons, including memory and consistent formatting, only court rulings from 2000 onward are part of the database.

I use a standard retrieve-rerank pipeline. A k-nearest neighbor search over the embedded documents and a BM25 keyword search over a trigram index (SQlite FTS5) each return their top-30 documents. BM25 should catch strict terminology, dense embeddings, on the other hand, the semantic meaning. Results are then reranked using a cross-encoder. All technical details, benchmarks and analyses are discussed and developed in separate posts but here is a high-level view of the pipeline:

         ┌─→ Dense: ChromaDB (cosine) ───→ top-30 ┐
Query ───┤                                        ├─→ merge ─→ cross-encoder ─→ top-n
         └─→ BM25: SQLite FTS5 (trigram) → top-30 ┘

Court rulings from 2000-2025 were scraped from the official website of the Swiss Federal Supreme Court and parsed using BeautifulSoup. In total the corpus spans 6'682 rulings in three of the national languages. The corpus is predominantly in German (66%) but also includes French (30%) and Italian (4%). All court rulings along with metadata were stored in a local sqlite database. The rulings were split by paragraph, which keeps them within the limits of the wide-spread 512 token limit of many embedding models:

Total paragraphs:122'398
Median:191 tokens
90th:440 tokens
95th:538 tokens
99th:751 tokens
Max:2'065 tokens

About 6% of paragraphs exceeded 512 tokens in length. For models with a 512 token window, these paragraphs were truncated.

At the ruling level, all cross-references to other rulings and legal articles are parsed as metadata. Secondly, also on the paragraph level, all norm references, ruling cross-references, literature citations and its consideration number were extracted. Each paragraph is hence identified by the tuple (ruling_id, erw_nr, para_nr).

Each paragraph was subsequently embedded using a sequence embedding model (see Part 2 for more details and benchmarking) and stored in a chromadb database for efficient retrieval.

Challenges

Even before discussing any precise findings and evaluations, two overarching challenges can be anticipated for the retrieval models.

Distributional Shift

Firstly, one needs to consider the distributional shift between the training corpus of the pre-trained embedding models and cross-encoders and of the Swiss legal domain. On the one hand, we require a model that handles multi-lingual input (German, French and Italian). On the other hand, Swiss “legalese” does not form a large part of the training corpus of these models. Given the density of domain-specific language in court rulings, this is a big distributional shift that forms a barrier for general-purpose models. The effect of domain-specific pretraining is analyzed in Part 3. Distributional Shift

Lacking Ground Truth Signal

The second challenge in designing an information retrieval system for the legal context lies in the ground truth signal. Mapping a query to the relevant legal provisions is already difficult for broad queries, but at least each legal aspect generally corresponds to identifiable articles. For rulings, no such mapping exists, so any ground truth is noisy. Citing a court ruling may be a first-match or most recent-match effort by the author, while other authors may consider a best match or the oldest ruling.

This problem is partly alleviated by the fact that BGE rulings are officially published as “landmark” rulings. They address a given legal aspect extensively, either for the first time or in a change of practice to keep the ruling collection condensed. Still a significant overlap persists since there are many recurring questions. In contrast to case law jurisdictions, where many legal questions have well-known authoritative cases, in Swiss law, such mappings are therefore much more noisy.

With these limitations in mind, a proxy for the relevance ground-truth signal can be extracted from legal commentaries (see Legal Materials above), where academics and practitioners treat legal provisions individually, discussing and summarizing their interpretation. Specifically, when a commentary paragraph cites a ruling by the Federal Supreme Court, it is generally to support the content in the paragraph. It is best-practice in the legal field to cite at a fine-grained level, wherefore generally one can expect a good restriction of the scope for a citation. Importantly, given the expertise of the authors, such paragraph-citation matchings include an expert-judgement on the most relevant citation that support a claim.

The “Online-Kommentar” is an open-access commentary for Swiss law that contains commentaries for around 180 legal provisions from all major areas of Swiss law. It is an ongoing project and accessible at https://onlinekommentar.ch. While it does not provide a complete coverage and the extensiveness of commentaries varies somewhat between different authors, it forms a good basis for this evaluation set.

In Part 2, I use these commentary citations to build a benchmark and evaluate how well pretrained embedding models retrieve the cited rulings.