100% Local  ·  RAG  ·  Laya Context Intelligence  ·  Expert Skills  ·  GGUF via llama.cpp

The Local Mind
for Your Circuits

ElectroMind AI bundles a llama.cpp inference server, a PyQt6 chat GUI and a retrieval-augmented engine that indexes your project — firmware, schematic PDFs, datasheets and docs — so every answer is grounded in your files, with citations.

🔒  No cloud. No API keys. No telemetry. Your code and chats never leave your machine.

Start in three steps See how it works

Windows x64  ·  Qwen3, Gemma 3 or any .gguf  ·  GPU optional

3 steps
install, drop a model, run — no accounts, no cloud setup
0 bytes
of your code, schematics or chats ever sent to any server
2 modes
⚡ Fast hash embeddings — instant & offline; 🎯 Accurate MiniLM — semantic
∞ chats
named, auto-titled and saved per project; long chats trimmed, never lost

What you get

One desktop app that keeps the model, the knowledge and the conversation on your side of the network.

Chat

Streaming chat with a thinking bubble

Answers stream into chat bubbles; reasoning models like Qwen3 get a separate thinking bubble so you can watch the reasoning. Multiple named chats per project, auto-titled from your first question and clickable to reopen — full history stays on disk even when long chats are trimmed before sending.

  • Editable system prompt, pre-tuned for electronics expertise
  • Context-window memory management for very long conversations
chat · ElectroMind
// you ask
Why does my buck converter ring at SW node?

// ElectroMind answers, citing your files:
┃ hardware/power/buck.pdf · hw/stm32f103.c
The ring is snubber resonance. Add RC across the
low-side FET: 4.7Ω + 1nF, keep traces < 10 mm…
[thinking] Switch-node ringing → parasitic L·C…
Models

Local GGUF models, one click away

Pick a model from the ./Model folder, or download one straight from the built-in catalog with a live progress bar — Qwen3, Gemma 3, or any .gguf you place there yourself.

  • Server settings: port, context size, GPU layers, temperature, max tokens
  • One-click start/stop with health checks; orphaned servers cleaned up automatically
Models panel
Qwen3 4B (Q4_K_M)      ~2.5 GB   ● Ready
Gemma 3 4B IT          ~2.5 GB   ○
Qwen3 8B (Q4_K_M)      ~5.0 GB   ○

context  8192   gpu layers  0
temp     0.7    max tokens  -1
▶ Start server   ■ Stop
Project Knowledge · RAG

It has actually read your files

Point ElectroMind at your project folder and hit Analyze. It indexes source code, configs, docs, PDF (layout mode tuned for CAD/schematic exports) and DOCX — then answers cite the exact files it used.

  • Incremental indexing — only new/changed files are re-processed
  • Fast mode needs zero downloads; Accurate mode uses MiniLM embeddings
indexing · my-project/
🔍 Analyze…
✔ 38 files scanned
✔ 6 PDFs extracted (layout mode)
✔ 214 chunks embedded → chroma_db/
✔ done in 3.2 s (incremental: 0 changed)
Laya Context Intelligence

A decision model filters your context

ChromaDB retrieves more candidates than needed — then Laya, a small local decision model, answers one question per chunk: "is this relevant to the user's question?" Sub-threshold chunks are dropped, the rest are re-ranked, and only the best few reach the LLM. One forward pass per chunk, ~100% local, and it never generates a word — it only decides.

  • Multilingual: auto-routes English, Persian and mixed-language questions
  • Optional — falls back to plain ChromaDB ranking when not installed
  • Relevance ranking, not fact-checking: a high score means topically relevant
chat · context selection
ChromaDB: 10 candidate chunks
🧠 Laya evaluated 10 contexts, selected 3
0.94  hardware/power/buck.pdf
0.89  firmware/stm32f103.c
0.83  docs/psu-design.md
0.12  unrelated.md   ← dropped
→ only the top 3 go to llama.cpp

Skills — pick the expert you need

Tick one or more expert roles in Settings ▸ Skills and their instructions are injected into the system prompt of every message. Combine them freely — PCB + STM32 + Power — no prompt engineering required.

⚡ Hardware & Power

Schematic & Circuit Designer · Analog & Signal-Conditioning · Power Supply · Battery & BMS · PCB Designer · Motor & Motion Control · Sensor & Instrumentation

📶 RF, Digital & Control

RF & Antenna Designer · FPGA / RTL (Verilog–VHDL) · Signal Processing / DSP · Control Systems

🔌 Firmware & Microcontrollers

Embedded Developer · STM32 · ESP32 / Arduino · AVR / PIC · RTOS / FreeRTOS · Bus & Protocol · Linux & Driver · IoT & Cloud

🐍 Software

Python Developer · C / C++ Software Engineer

🛡️ Quality, Safety & Debugging

Debugging & Test · Test & Validation · Safety & Security · PLC & Industrial Automation · Robotics

⭐ Your own skills

Every skill is a Markdown file in ./skills. Create skills in the dialog, edit the files, or drop new .md files in the folder — they appear automatically, no restart needed.

Context-aware by design: ElectroMind estimates the prompt size of the ticked skills and auto-sizes the model's context window so everything fits. If a request would still overflow, the system message is trimmed — project context first, then skills, then prompt — and the app tells you exactly what was cut. With Laya Context Intelligence enabled, the project context itself is re-ranked by a local decision model before it reaches the LLM.

Settings ▸ Skills · ElectroMind
Available skills

HARDWARE & POWER
☑ ⚡ Schematic & Circuit Designer
☑ 🔋 Power Supply Designer
☐ 🖨️ PCB Designer

FIRMWARE & MICROCONTROLLERS
☑ 🔧 STM32 Developer
☐ 📡 ESP32 / Arduino Developer

Active: ⚡ Schematic · 🔋 Power · 🔧 STM32
skills are added to the system prompt for every message

How it works

Four stages, from your file system to a cited answer. Everything runs as local processes on your PC.

01

Scan

Walks your project folder and skips noise — .git, build, node_modules… Files over 5 MB are reported and skipped.

02

Load & chunk

Extracts text — PDF via pypdf layout mode with CAD-export cleanup, DOCX including tables — then splits into ~1500-char overlapping chunks.

03

Embed & store

Chunks go into ChromaDB with their relative path as metadata, using fast hash vectors or MiniLM embeddings — both local.

04

Retrieve & answer

Each question pulls the top-4 most similar chunks into the prompt as project context; the answer appears with its source files underneath.

Built in three clean layers

ui/

PyQt6 dark-themed desktop GUI: chat, models, RAG and a tabbed settings window (server, models, RAG, skills, system prompt), with background worker threads so the interface never freezes.

core/

GUI-independent logic: llama-server process control, proxy-safe localhost HTTP, per-project chat persistence, the model catalog and the skills engine that parses ./skills/*.md into prompt-ready expert roles.

rag/

The retrieval pipeline: file discovery, text/PDF/DOCX extraction, chunking, embeddings and the ChromaDB store.

Getting started

Python 3.10+ on Windows x64. Three steps and you're chatting with your own hardware docs.

1

Install dependencies

One pip command pulls PyQt6, ChromaDB and the document extractors:

pip install -r requirements.txt
2

Get a model

Drop any .gguf into ./Model — or download one from inside the app's catalog. Qwen3 4B (Q4_K_M) (~2.5 GB) is the recommended start: good CPU speed/quality balance. Have a GPU-enabled llama.cpp build? Set GPU layers > 0.

python gui.py        # or: python run_llama.py (terminal-only chat)
3

Connect your project & ask

  1. Select a model → ▶ Start server → wait for ● Ready
  2. 📂 Browse your project folder → 🔍 Analyze to index it
  3. ⚙ Settings ▸ Skills — tick the expert roles you need (optional, but this is where the magic is)
  4. Ask questions — answers arrive grounded in your files, with citations

Every answer still needs engineering judgement: the model is text-only (no vision), and critical values — mains voltage, Li-ion/LiPo handling, ESD, current limits — must always be verified against the datasheet. The built-in system prompt enforces safety-aware answers, but you are the engineer.

FAQ

The questions that actually come up.

Does anything I write get sent to the internet?⌄

No. Everything runs locally — the llama.cpp server, the embeddings, the vector store. Only model downloads (the GGUF file and, optionally, the MiniLM embedding model) touch the internet. Localhost requests even bypass the Windows system proxy.

Can it read my Altium/KiCad schematic files?⌄

Native CAD binaries (like .SchDoc) have no text layer, so they're skipped — export them as PDF and re-Analyze. The PDF loader is tuned for CAD/schematic exports and reads the text layout, not images. Scanned-image PDFs likewise contain no text to read.

The app closes instantly after Analyze — why?⌄

A known Windows DLL conflict (onnxruntime vs Qt) that is fixed by design: gui.py imports onnxruntime before Qt. If you ever see it again, delete chroma_db\ef_name.txt and restart.

Port 8081 is already in use — what now?⌄

Pick another port in the settings, or stop the other server. The app also cleans up leftover llama-server.exe processes on start.

Which model should I start with?⌄

Qwen3 4B (Q4_K_M) — about 2.5 GB, a good balance of quality and speed on CPU for chat, code and electronics Q&A. Need it smarter? Qwen3 8B needs ~8 GB RAM and more patience. Just testing? Qwen3 0.6B runs on almost anything.

Can it view images or photos of my board?⌄

No — local GGUF text models have no vision. Describe the image in your question, or export the schematic as a text-based PDF and let RAG read it.

What exactly do Skills do — and can I add my own?⌄

Skills are selectable expert roles whose instructions are added to the system prompt for every message — tick STM32 Developer and answers about your firmware follow CubeMX/HAL conventions and errata-aware patterns. 26 are built in; your own are just Markdown files in ./skills (name, description, ## Instructions) created in the dialog or dropped in by hand — they appear automatically. Ticked skills are auto-sized into the context window, and the app tells you when anything had to be trimmed.

Do skills and RAG work together?⌄

Yes — that's the intended combo: your indexed project files provide the facts, the skills shape the expertise of the answer. The context auto-sizer accounts for both, so adding skills won't silently crowd out your project context.

What is Laya Context Intelligence?⌄

An optional extra layer: after ChromaDB retrieves candidate chunks, a small local decision model (Laya, 322–421 M parameters) scores each chunk's relevance to your question in a single forward pass. Irrelevant chunks are dropped before the LLM ever sees them, so answers lean on fewer, better files. It's relevance selection, not fact-checking — and if Laya isn't installed, ElectroMind simply uses the plain ChromaDB ranking. Enable it in Settings ▸ Project / RAG (pip install laya, Python 3.10+).