100% Local · RAG · Laya Context Intelligence · Expert Skills · GGUF via llama.cpp
ElectroMind AI bundles a llama.cpp inference server, a PyQt6 chat GUI and a retrieval-augmented engine that indexes your project — firmware, schematic PDFs, datasheets and docs — so every answer is grounded in your files, with citations.
🔒 No cloud. No API keys. No telemetry. Your code and chats never leave your machine.
One desktop app that keeps the model, the knowledge and the conversation on your side of the network.
Answers stream into chat bubbles; reasoning models like Qwen3 get a separate thinking bubble so you can watch the reasoning. Multiple named chats per project, auto-titled from your first question and clickable to reopen — full history stays on disk even when long chats are trimmed before sending.
// you ask
Why does my buck converter ring at SW node?
// ElectroMind answers, citing your files:
┃ hardware/power/buck.pdf · hw/stm32f103.c
The ring is snubber resonance. Add RC across the
low-side FET: 4.7Ω + 1nF, keep traces < 10 mm…
[thinking] Switch-node ringing → parasitic L·C…
Pick a model from the ./Model folder, or download one straight from the built-in
catalog with a live progress bar — Qwen3, Gemma 3, or any .gguf you place there
yourself.
Qwen3 4B (Q4_K_M) ~2.5 GB ● Ready
Gemma 3 4B IT ~2.5 GB ○
Qwen3 8B (Q4_K_M) ~5.0 GB ○
context 8192 gpu layers 0
temp 0.7 max tokens -1
▶ Start server ■ Stop
Point ElectroMind at your project folder and hit Analyze. It indexes source code, configs, docs, PDF (layout mode tuned for CAD/schematic exports) and DOCX — then answers cite the exact files it used.
🔍 Analyze…
✔ 38 files scanned
✔ 6 PDFs extracted (layout mode)
✔ 214 chunks embedded → chroma_db/
✔ done in 3.2 s (incremental: 0 changed)
ChromaDB retrieves more candidates than needed — then Laya, a small local decision model, answers one question per chunk: "is this relevant to the user's question?" Sub-threshold chunks are dropped, the rest are re-ranked, and only the best few reach the LLM. One forward pass per chunk, ~100% local, and it never generates a word — it only decides.
ChromaDB: 10 candidate chunks
🧠 Laya evaluated 10 contexts, selected 3
0.94 hardware/power/buck.pdf
0.89 firmware/stm32f103.c
0.83 docs/psu-design.md
0.12 unrelated.md ← dropped
→ only the top 3 go to llama.cpp
Tick one or more expert roles in Settings ▸ Skills and their instructions are injected into the system prompt of every message. Combine them freely — PCB + STM32 + Power — no prompt engineering required.
Schematic & Circuit Designer · Analog & Signal-Conditioning · Power Supply · Battery & BMS · PCB Designer · Motor & Motion Control · Sensor & Instrumentation
RF & Antenna Designer · FPGA / RTL (Verilog–VHDL) · Signal Processing / DSP · Control Systems
Embedded Developer · STM32 · ESP32 / Arduino · AVR / PIC · RTOS / FreeRTOS · Bus & Protocol · Linux & Driver · IoT & Cloud
Python Developer · C / C++ Software Engineer
Debugging & Test · Test & Validation · Safety & Security · PLC & Industrial Automation · Robotics
Every skill is a Markdown file in ./skills. Create skills in the dialog, edit
the files, or drop new .md files in the folder — they appear automatically,
no restart needed.
Context-aware by design: ElectroMind estimates the prompt size of the ticked skills and auto-sizes the model's context window so everything fits. If a request would still overflow, the system message is trimmed — project context first, then skills, then prompt — and the app tells you exactly what was cut. With Laya Context Intelligence enabled, the project context itself is re-ranked by a local decision model before it reaches the LLM.
Available skills
HARDWARE & POWER
☑ ⚡ Schematic & Circuit Designer
☑ 🔋 Power Supply Designer
☐ 🖨️ PCB Designer
FIRMWARE & MICROCONTROLLERS
☑ 🔧 STM32 Developer
☐ 📡 ESP32 / Arduino Developer
Active: ⚡ Schematic · 🔋 Power · 🔧 STM32
skills are added to the system prompt for every message
Four stages, from your file system to a cited answer. Everything runs as local processes on your PC.
Walks your project folder and skips noise — .git, build,
node_modules… Files over 5 MB are reported and skipped.
Extracts text — PDF via pypdf layout mode with CAD-export cleanup, DOCX including
tables — then splits into ~1500-char overlapping chunks.
Chunks go into ChromaDB with their relative path as metadata, using fast hash vectors or MiniLM embeddings — both local.
Each question pulls the top-4 most similar chunks into the prompt as project context; the answer appears with its source files underneath.
ui/PyQt6 dark-themed desktop GUI: chat, models, RAG and a tabbed settings window (server, models, RAG, skills, system prompt), with background worker threads so the interface never freezes.
core/GUI-independent logic: llama-server process control, proxy-safe localhost HTTP, per-project chat
persistence, the model catalog and the skills engine that parses ./skills/*.md
into prompt-ready expert roles.
rag/The retrieval pipeline: file discovery, text/PDF/DOCX extraction, chunking, embeddings and the ChromaDB store.
Python 3.10+ on Windows x64. Three steps and you're chatting with your own hardware docs.
One pip command pulls PyQt6, ChromaDB and the document extractors:
pip install -r requirements.txt
Drop any .gguf into ./Model — or download one from inside the app's
catalog. Qwen3 4B (Q4_K_M) (~2.5 GB) is the recommended start: good CPU
speed/quality balance. Have a GPU-enabled llama.cpp build? Set GPU layers > 0.
python gui.py # or: python run_llama.py (terminal-only chat)
Every answer still needs engineering judgement: the model is text-only (no vision), and critical values — mains voltage, Li-ion/LiPo handling, ESD, current limits — must always be verified against the datasheet. The built-in system prompt enforces safety-aware answers, but you are the engineer.
The questions that actually come up.
No. Everything runs locally — the llama.cpp server, the embeddings, the vector store. Only model downloads (the GGUF file and, optionally, the MiniLM embedding model) touch the internet. Localhost requests even bypass the Windows system proxy.
Native CAD binaries (like .SchDoc) have no text layer, so they're skipped — export
them as PDF and re-Analyze. The PDF loader is tuned for CAD/schematic exports and reads the text
layout, not images. Scanned-image PDFs likewise contain no text to read.
A known Windows DLL conflict (onnxruntime vs Qt) that is fixed by design: gui.py
imports onnxruntime before Qt. If you ever see it again, delete
chroma_db\ef_name.txt and restart.
Pick another port in the settings, or stop the other server. The app also cleans up leftover
llama-server.exe processes on start.
Qwen3 4B (Q4_K_M) — about 2.5 GB, a good balance of quality and speed on CPU for chat, code and electronics Q&A. Need it smarter? Qwen3 8B needs ~8 GB RAM and more patience. Just testing? Qwen3 0.6B runs on almost anything.
No — local GGUF text models have no vision. Describe the image in your question, or export the schematic as a text-based PDF and let RAG read it.
Skills are selectable expert roles whose instructions are added to the system prompt for every
message — tick STM32 Developer and answers about your firmware follow CubeMX/HAL
conventions and errata-aware patterns. 26 are built in; your own are just Markdown files in
./skills (name, description, ## Instructions) created in the dialog or
dropped in by hand — they appear automatically. Ticked skills are auto-sized into the context
window, and the app tells you when anything had to be trimmed.
Yes — that's the intended combo: your indexed project files provide the facts, the skills shape the expertise of the answer. The context auto-sizer accounts for both, so adding skills won't silently crowd out your project context.
An optional extra layer: after ChromaDB retrieves candidate chunks, a small local decision model
(Laya, 322–421 M parameters) scores each chunk's relevance to your question in a single
forward pass. Irrelevant chunks are dropped before the LLM ever sees them, so answers lean on
fewer, better files. It's relevance selection, not fact-checking — and if Laya isn't installed,
ElectroMind simply uses the plain ChromaDB ranking. Enable it in Settings ▸ Project / RAG
(pip install laya, Python 3.10+).