An AI assistant for my Lexus CT200h's workshop manual

A RAG system that answers questions about a car's workshop manual and cites its pages: hybrid search, diagnostic trees, wiring connections rebuilt from the diagrams, and web search for known faults.

RAGPythonLanceDBLLM
Contents
  1. What a RAG is, briefly
  2. The material
  3. Why semantic search gets codes wrong
  4. Reranking the candidates
  5. The manual is not flat text
  6. Rebuilding connections from the wiring diagrams
  7. Reading, not just retrieving
  8. The manual isn't enough: web search
  9. Problems and fixes
  10. How I check changes
  11. Limits
  12. Stack

In short

  • I built an assistant that answers questions about the workshop manual of my Lexus CT200h: about 13,500 files in English, including HTML pages, wiring diagrams and PDFs.
  • A language model on its own is not enough: you first have to find it the right pages. That is the hard part, not generating the text.
  • Codes (fault codes, connectors, pins) have to be searched by exact match, descriptions by meaning. You need both.
  • In the wiring diagrams the connections are not written down anywhere, so I rebuilt them from the geometry of the drawing.
  • The manual says how to check a fault, not which faults actually happen: that takes web search, over service bulletins and forums.
  • Every claim in the answer cites the page it comes from, manual or web, so it can be checked.

I have the workshop manual and the wiring diagrams for my Lexus CT200h, a compact hybrid from Toyota's premium brand. Looking things up by hand is slow, and my questions are rarely "what's on page 412". They are more like "I have code P0A80, where do I start?" or "what is pin 42 of the ABS ECU (the braking system's control unit) connected to, and what voltage should I measure there?".

CarManualRAG is the prototype I wrote to answer questions like these. It runs locally, in Python. For generation it uses any OpenAI-compatible service: a local model in LM Studio (an app for running language models on your own computer) or a cloud provider, with no changes to the rest of the system.

What a RAG is, briefly

RAG stands for retrieval-augmented generation: first you find the relevant documents, then you pass them to the language model together with the question, and the model answers from them. A language model on its own knows a lot about cars in general, but it doesn't know the expected voltage on a specific pin of my ABS ECU. Without the right page, it makes up a plausible number.

An analogy. It's the difference between asking an experienced mechanic to answer from memory and asking them to answer with the manual open in front of them. RAG is the part that opens the manual at the right page. If the page is wrong, even the best mechanic gets it wrong.

Almost all the work in this project is in the retrieval, not the generation.

The material

The manual comes as a static HTML site: 5,480 pages, some of which contain SVG wiring diagrams, plus 24 PDFs and 8,020 images. It is all in English, while I ask my questions in Italian.

The parser splits the pages into text blocks of at most 1,800 characters (chunks) and assigns each one a document type. The final index holds 28,114 chunks:

Page typeChunksWhat it contains
Fault code (DTC) procedures10,081what to do, step by step
General4,951pages that fit none of the other types
Connectors4,186connector shape, pin numbering
Repair3,698removal, installation, torque values
Circuit inspection1,233checks on a single circuit
Symptom tables936symptom → suspect areas, most likely first
System descriptions832how a system works
Data lists682values read with the scan tool
Control unit (ECU) terminals452ECU pinouts with expected values
Components, DTC charts, specs, schematics, parts location, PDF1,063the rest

The type is used to narrow the search: if the question is about a pin, there is no point searching the removal procedures.

Why semantic search gets codes wrong

The first attempt is the obvious one: compute an embedding for each chunk, a vector of numbers that represents its meaning, and look for the chunks whose vectors are closest to the question's. I use BGE-M3, a multilingual embedding model.

For descriptive questions ("the radio doesn't detect the parking brake") it works well. It fails exactly where precision matters. To an embedding model P0A80 and P0A8F are nearly the same thing, and a connector like H164-15 or a signal like CANH is little more than noise. In a workshop manual, the difference between two codes is everything.

So the search is hybrid. Next to the vector, each chunk has a text field with the codes it contains: fault codes in the standard on-board diagnostics (OBD) form (one of P, B, C, U followed by four characters), connectors with pin numbers, signal names. I extract them with deliberately conservative rules: I'd rather miss a few codes than fill the index with false positives. That field is indexed with BM25, the classic text-search ranking algorithm, which rewards exact matches and rare terms.

In practice. Search by meaning finds "the page about noises from the roof" even when the words differ. Search by exact term finds "the page containing P0A80" and nothing else. The two result lists are merged.

LanceDB does both in the same file-based database, with no server to run: vectors on one side, a Tantivy full-text index on the other.

Reranking the candidates

Hybrid search is fast but coarse. For each question it returns about thirty candidates, and the first one is not always the best. A second model, a cross-encoder (bge-reranker-v2-m3), reads each candidate together with the question and gives it a relevance score between 0 and 1. I keep the top eight.

A cross-encoder is far more accurate than vector search because it looks at question and document together, instead of comparing two vectors computed separately. It is also far slower, which is why it only runs on a few candidates and not on the whole index.

Question in Italian "tettuccio in vetro che vibra" Rewrite into English keywords the original question is kept too Hybrid search: meaning + codes about thirty candidates per search, then merged Reranker (cross-encoder) score 0–1, the top 8 are kept Agent with tools reads pages, follows procedures, searches the web Answer with citations every fact links back to the manual page
The path of a question. The language model only comes in at the rewrite and in the last two stages.

The manual is not flat text

Diagnostic procedures are trees. Each step says what to measure and has two outcomes: OK (within spec) and NG (no good, out of spec). Each outcome leads to a different step.

1 · Measure voltage at the pin OK NG 2 · Check the sensor 4 · Check the harness
A simplified example. Real procedures can run to ten or fifteen steps.

If you flatten a procedure into paragraphs and split it into chunks, the model sees fragments of both branches and ends up mixing them: it suggests replacing the sensor when the measurement said the problem was the wire. The parser instead stores the structure of the tree, with the OK and NG links of every step. The model can ask for the procedure of a code and then advance one step at a time, stating which outcome it got.

In the same way, ECU terminal tables are reassembled: a long table ends up split across twenty or so chunks, and answering a question about a pin needs the right row together with its headers and units.

Rebuilding connections from the wiring diagrams

In the wiring diagrams, connectivity is not written anywhere. In the SVG file, connectors are rectangles, pins are coordinates, wires are coloured polylines and labels are loose text scattered over the sheet. A person looks at the drawing and immediately sees where a wire goes. A system that only reads text cannot answer "what is this pin connected to?", which is half of electrical diagnosis.

The reconstruction is purely geometric. For each wire the parser takes its two endpoints and looks for the nearest connector, measuring the distance to its rectangle. Then it looks for text labels within a small radius of each endpoint: the pin number, the wire colour code in Toyota's convention and the signal name. The result is a netlist, the list of connections pin by pin.

A61 ABS ECU A29 pedal stroke sensor 42 SKS R 4 SKS1 dashed circles: radius searched for labels
From drawing to list: A61 pin 42 (SKS) → A29 pin 4 (SKS1), red wire. The parser only sees rectangles, lines and text.

It works for wires that end directly on a connector, which is between 45 and 60% of the wires in each diagram. It doesn't yet follow intermediate splices, where a wire branches into several.

The most useful check was cross-referencing two independent sources. The ABS ECU terminal table says pin 42 is SKS, the brake pedal stroke sensor input, with its expected value. The netlist rebuilt from the diagram says pin 42 goes to connector A29, which is the pedal stroke sensor. The two facts come from different pages and different parsers, and they agree. When the agent uses them together it can give a measurable instruction: which pin, to which connector, what wire colour, what reading to expect.

Reading, not just retrieving

Classic RAG puts the top eight chunks in the prompt and asks the model to answer. Instead, I give the model a list of candidate documents and a set of tools, and let it decide what to read.

ToolWhat it does
Searchsearches the index, optionally within one page type
Readreads a whole page or diagram
DTC procedurereturns the diagnostic tree for a fault code
Next stepadvances in the tree given an OK or NG outcome
Connectorreturns a connector's pin-to-pin connections
ECU terminalsreturns the expected value at an ECU pin
Symptomsuspect areas from the symptom table, in order
Imageshows the rendered diagram to models that can see images
Web searchtechnical bulletins and known issues outside the manual

The reason is simple: a small chunk is great for finding the right page, but it often lacks a procedure's preconditions or the measurement conditions, which sit at the top of the page. Reading the full page brings that context back.

Every claim in the answer carries the identifier of the document it comes from. In the web interface, clicking a citation opens the original page next to the chat, with the cited passage highlighted. Web sources get the same kind of citation. This is what makes the tool usable: I don't have to trust the answer, I can check it in two seconds.

The agent has a budget of 40 steps per question. I tried Google's Gemma locally in LM Studio, and DeepSeek V4 Flash and Xiaomi's MiMo v2.5 through their APIs. The backend is the same for all of them; what changes is the quality of the reasoning, how willing they are to use tools and whether they can look at diagrams as images.

The workshop manual describes how the car should work and how each component is checked. It doesn't say which faults actually happen, how often, or at what mileage. That information lives elsewhere: in the manufacturer's technical service bulletins (TSBs), owners' forums, mechanics' discussions. For a diagnosis it's often the most useful part, because it tells you where to start.

So the agent also has a web search tool. It uses DuckDuckGo, so no API key is needed, and returns the top five results. In the prompt I split the roles explicitly:

SourceWhat you find thereExample question
Manual (local index)specs, procedures, expected values, diagrams"what voltage should I measure at pin 42?"
Webrecurring faults, service bulletins, real-world experience"is this symptom a known issue on this engine?"

The two sources complement each other. The web suggests the most likely hypothesis; the manual says how to confirm it with a measurement. A good answer uses both: "it's a known issue, here's the manual procedure to confirm it".

The example that convinced me is the engine vibrating when cold. The manual has every procedure for checking ignition, injection and engine mounts, but no page says that on this engine the head gasket is a known weak point. On the web, it's widely discussed. Once the model started using web search it reached that hypothesis, and from there you move on to the manual's checks.

An analogy. The manual is the instruction book; the web is the old mechanic who says "on that model, it's always this". You need both: the first tells you how to check, the second where to start.

The web is a less reliable source than the manual, though, so it has to stay recognisable. The model must say where each piece of information comes from: manual, web or general experience. Web sources get their own numbered citations, just like manual pages. Many sites refuse to be shown inside another page, so clicking a web citation opens a preview card with the title, an excerpt and a link to the original.

There was an instructive bug here too. Web source numbering restarted at 1 with every message: on the second turn "web2" pointed to a different page than on the first, and citations in earlier messages led to the wrong source. Numbering now continues across the whole conversation.

Problems and fixes

The problems that came up during development are the most useful part to tell, because they show up in almost every RAG system.

The model that answers from memory

With small models the first problem was that they didn't use the tools. The logs showed the model "planning" to look up the symptom table and then answering from memory, based on the candidate snippets already in front of it. On a question about engine vibration when cold, it never got to the known head gasket issue, which is well documented online.

The fix had two parts. In the prompt I described what is in the local index (exact specs and procedures) and what is on the web (real-world defects and bulletins), because the model didn't know what it was querying. And if the first answer comes back with no tool calls at all, it is discarded and the model is asked, once, to check the sources.

That fix introduced a bug of its own. At first the check ran on every message, not just the first. So on the second or third turn, when the model legitimately answered without tools (to rule out a hypothesis or ask a question, say), the answer vanished halfway through. Now it only applies to the first answer of a conversation.

Off-topic candidates

The manual is in English. A vague question in Italian, searched as is, returned pages with no connection to it, about power steering or the immobiliser for example, and small models got dragged along. The reranker, however, told them apart well: off-topic pages scored close to zero, relevant ones between 0.5 and 0.99. Now the candidates shown to the model must clear a threshold of 0.05. If none does, the model gets no list at all and is told to search with the tools, instead of starting from a misleading list.

The sunroof that turned into the hybrid system

Before searching, a call to the model rewrites the question into English keywords, leaving codes untouched. At one point the question "tettuccio in vetro che vibra" (glass roof rattling) returned the hybrid system's Loud Rattle page.

The reranker wasn't to blame: it gave 0.90 to a relevant page and 0.00 to noise. The problem was earlier. The rewrite produced sunroof, a word that never appears in the manual (it says sliding roof), and added Lexus CT200h rattle. Since I only searched with the rewritten question, the roof pages never even made it into the candidates, so the reranker had nothing to recover. The original Italian question, on the other hand, found them first, because BGE-M3 is multilingual.

The fix has two parts. I now search with both the rewritten and the original question, merge the two candidate sets and rerank the union. In the rewrite prompt I added the manual's own vocabulary with common synonyms ("tettuccio" becomes sliding roof sunroof moonroof) and removed make and model, which narrow nothing in a corpus about a single car.

A fixed list of words to drop doesn't work: "F-Sport" (the sporty trim level) is noise in a question about the roof but essential in one about the suspension; "LED" is noise in the first and central to a question about headlights. Now relevance is decided by the rewrite, question by question.

The lesson applies beyond this project: if a document doesn't make it into the candidates, no reranker can recover it. When the result is wrong, check recall (how many right documents reach the reranking stage) before touching the reranking.

The prompt that saw head gaskets everywhere

The cold-vibration case, used as a test, had ended up in the system prompt as an example of reasoning: cold, therefore head gasket or condensation. An example that specific steers the model more than any general instruction: every diagnosis leaned towards the engine, even ones about the sunroof, the brakes or the air conditioning.

I replaced the example with a general principle: let the conditions (cold or hot, over bumps, under load, in the rain, gradual or sudden, with or without a fault code) point to the kind of fault. Then I re-ran questions about different systems to check that each stayed within its own area.

It's an easy mistake to make with prompts: you fix one case and make all the others worse, with nothing to warn you.

How I check changes

Checks are based on a set of reference questions that already worked. After every change to the retrieval I check that the right document stays on top, and with what reranker score:

Reference questionScore of the top document
Code P0A800.86
Pin 42 of ABS ECU A610.96
Daytime running lights (DRL)0.84
The radio doesn't detect the parking brake0.31

The last score is lower because the question is descriptive and the page uses different words. The absolute value doesn't matter: what matters is that these numbers don't get worse from one version to the next. For the agent's full answers the check is still manual: I read the answer and open the citations.

Limits

  • The netlist rebuilds only part of the connections.
  • There is no systematic evaluation: I have regression questions, not a benchmark.
  • PDFs that are mostly drawings need a model that can see images.
  • Sessions live in server memory; conversation history is stored in the browser.

It's a prototype for understanding how to search complex technical documentation, not a tool to diagnose a car instead of a workshop. The code and the manuals are not public.

Stack

PartChoice
EmbeddingsBGE-M3 via sentence-transformers (MPS on Apple Silicon)
Rerankingbge-reranker-v2-m3 (cross-encoder)
IndexLanceDB: vectors + Tantivy full-text, file-based
Generationany OpenAI-compatible endpoint (LM Studio or cloud)
ParsingBeautifulSoup, PyMuPDF, cairosvg for SVG rendering
InterfaceFastAPI + build-free HTML/JS, streamed answers (SSE)

← All posts