The people who most need private AI
are the least equipped to build it.
Law firms, RIAs, medical practices, and accounting firms hold data they cannot lawfully or commercially place in a public cloud: case files, health records, portfolio allocations. The market forces them to choose between easy and private. Cloud AI is easy but not private. Enterprise on-prem hardware is private but not easy. Local-first tooling is a kit, not a product. Nobody sells private and turnkey as a single ordered object.
Silo Server closes that gap by turning private infrastructure into a consumer product. Each unit ships from the factory pre-configured with the full operational stack: local language models, retrieval over the firm's own documents, ingestion, access control, and the vertical template for the buyer's industry. The customer configures the machine online the way they would configure a Mac, and it arrives ready to run. Nothing phones home. The only network the machine knows is the one inside the building.
Capacity is physical. When the firm outgrows one unit, it orders another and clicks it onto the stack. The modules align magnetically, lock mechanically, and join a private fabric automatically. No rack, no integrator, no migration project. Scaling infrastructure becomes a purchasing decision instead of an engineering one.
This is not a cold start. Silo Systems has deployed this class of infrastructure by hand for consulting clients in regulated legal and financial services since 2019, under the internal designation Obox. The consulting practice proved the demand, funded the R&D, and supplies the design partners. Silo Server is the productization of work clients already pay for, aimed at the thousands of firms one consultant cannot reach.
The concept carried open engineering forks: compute platform, interconnect physics, memory pooling, OS, inference engine, data layer, model licensing, pricing, and manufacturing strategy. Every fork is closed in this revision with a reasoned decision, recorded in the Decision Register (§XI, D/01–D/12), each with the options weighed and the trigger that would reopen it.
A compliance budget
looking for a product.
The buyer is a managing partner or practice principal at a 1–50 seat regulated firm. They already spend on compliance, IT, and per-seat software. They have watched competitors adopt AI and they have read their own bar opinions, HIPAA guidance, and client engagement letters. Their constraint is not budget. It is that every AI product offered to them requires surrendering custody of the one asset they are professionally obligated to protect.
Why now, specifically. Three curves crossed between 2024 and 2026. Open-weight models reached genuine professional utility at the 30–70B scale. Unified-memory edge silicon made 70B-class local inference possible in a shoebox instead of a rack. And compliance pressure on cloud AI hardened from theoretical to explicit, in bar ethics opinions, insurer questionnaires, and client outside-counsel guidelines. The window where a local appliance is both feasible and differentiated is open now.
| Segment | US sites (est.) | Compliance driver | Beachhead priority |
|---|---|---|---|
| Law firms | ≈ 450,000 | ABA Model Rule 1.6 · privilege · OCG clauses | 01 · primary wedge |
| Registered investment advisers | ≈ 15,000 | GLBA · SEC 17a-4 · Reg S-P | 02 · fast follower |
| CPA & accounting firms | ≈ 46,000 | IRC §7216 · GLBA safeguards | 03 |
| Medical & dental practices | ≈ 230,000 | HIPAA · state privacy acts | 04 · post-v1 |
| Addressable regulated sites | ≈ 740,000 | SAM: ≈ 180,000 sites at 5–50 seats with an active compliance driver | |
| Path | Year 1 | Years 2–3 | 3-yr total | Custody of client data |
|---|---|---|---|---|
| Copilot-class cloud AI · $30/seat/mo | $7,200 | $14,400 | $21,600 | Transits vendor cloud · perpetual |
| Enterprise on-prem build · integrator + rack | $60,000+ | $24,000+ | $84,000+ | Retained · staff burden retained too |
| Silo Server Studio + Care | $8,487 | $5,976 | $14,463 | Never leaves the building |
Silo path: Studio $5,499 hardware + Silo Care $249/mo. The appliance is cheaper than cloud per-seat pricing by year two, and it is the only path where the answer to a client audit question is a photograph of a box in a locked room.
| Alternative | Private by physics | Turnkey | Scales by module | Zero-access service model |
|---|---|---|---|---|
| Cloud AI APIs · OpenAI, Anthropic, Azure | No | Yes | N/A | No |
| M365 Copilot · per-seat | No | Yes | N/A | No |
| Dell · HPE · NVIDIA on-prem | Yes | No · integrator required | Rack-scale only | No |
| AI workstations · Lambda-class | Yes | No · a computer, not a stack | No | No |
| DIY local · Ollama, Mac Studio | Yes | No · a hobby | No | No |
| Silo Server | Yes | Yes | Yes | Yes · §VI |
Hyperscalers cannot follow. Their economics require centralizing customer data; "it never leaves your building" is a promise their business model cannot make. Hardware incumbents can build boxes but sell through integrators and have no vertical software, no compliance narrative, and no appetite for 20-seat law firms. The moat is the combination: appliance + vertical stack + ordering experience + zero-access service model, aimed at a buyer everyone else finds too small to integrate and too regulated to cloud.
One silicon decision
carries the whole line.
The concept document left the compute platform open between Apple Silicon, NVIDIA, and x86. This revision closes it: AMD Strix Halo (Ryzen AI Max 300 series) across every tier, differentiated by unified memory capacity alone.
The reasoning is recorded in full at D/01, but the shape of it: Apple does not sell M-series silicon to third parties, so an Apple-based appliance cannot legally exist at scale. NVIDIA Jetson Thor is a strong inference module but binds the entire software estate to CUDA on ARM at a materially higher cost per gigabyte of unified memory. Strix Halo delivers up to 128GB of unified LPDDR5X, an XDNA NPU plus a 40-CU RDNA GPU, full x86 compatibility, and, decisively, a supply chain already proven open to small integrators. One architecture means one NixOS image, one thermal design, one test jig, and one spares bin across the entire product line. For a build-to-order operation, fleet homogeneity is not a preference. It is the support model.
The tier matrix
Three base units, one chassis, one image. Tiers are separated by unified memory and storage only, which is what actually gates local model capability. The concept document's TOPS figures are replaced with the platform's real envelope: a 50-TOPS XDNA NPU plus the RDNA GPU, with the GPU carrying LLM inference and the NPU carrying embedding and vision pipelines.
Model capability is stated in resident terms, not benchmark terms. A Pro holds a 70B model quantized to 4-bit, roughly 42GB, permanently resident with headroom for a concurrent 8B interactive lane, the vector index, and the application stack. That is the honest spec a buyer can hold us to.
Thermals, honestly
The concept called for fanless. Pure passive cooling caps sustained load near 25W, which cannot serve a 70B model, so the claim is revised to silent-first (D/02): a vapor chamber into a full-width fin stack, with two low-RPM assisted-convection fans that engage only under sustained inference. Mini runs passive in typical duty. The acoustic contract is printed on the spec sheet: under 24 dBA idle, under 32 dBA sustained at one meter. A machine sold to law offices is sold on silence, and we specify it like a dimension.
| Spec | Silo Server Mini | Silo Server Studio | Silo Server Pro |
|---|---|---|---|
| Seats served | 1–3 users | 5–15 users | 15–50 users |
| SoC | Ryzen AI Max · 12-core class | Ryzen AI Max · 16-core class | Ryzen AI Max+ 395 · 16-core |
| Unified memory | 32GB LPDDR5X-8000 | 64GB LPDDR5X-8000 | 128GB LPDDR5X-8000 |
| Storage · ZFS encrypted | 1TB NVMe | 2TB NVMe | 4TB NVMe |
| Resident models | ≤14B Q4 + embed + vision | 32B-class Q4 + 8B lane | 70B Q4 resident + 8B lane |
| Accelerators | XDNA NPU 50 TOPS (embed · OCR · vision) + RDNA GPU up to 40 CU (LLM inference) | ||
| Power · typ / peak | 65W / 120W | 90W / 160W | 110W / 240W |
| Acoustics @ 1m | ≤24 dBA · passive-biased | ≤28 dBA | ≤32 dBA sustained |
| I/O · security | 2×10GbE · 2×USB-C key-authenticated · TPM 2.0 · measured boot · internal 240W PSU · C14 | ||
| Chassis | 210 × 210 × 130 mm grid module · 6063-T5 + CNC · burnt-copper Type II anodize · 2.9 kg | ||
Manufacturing posture · integrate, don't fabricate
Version one is an integration business, not a fabrication business (D/01, manufacturing clause). Mainboards are sourced as Strix Halo OEM modules from the same supply chain that already serves small system builders. The chassis is extruded and CNC-finished at a job shop in lots of 50–200, anodized to the burnt-copper specification. Assembly, flashing, burn-in, and QA happen in-house in California, which is both a cost decision and a brand asset: every unit ships with a serialized certificate of provenance, assembled and attested where the company lives. Contract manufacturing is milestone-gated at a sustained 500 units per year.
Regulatory engineering: internal power is a pre-certified PSU module, which confines compliance work to FCC Part 15 Class B emissions testing and ETL listing of the assembly. Budgeted at $30k and eight weeks inside Phase 2 of the roadmap (§X).
Scaling you can do
with your hands.
The signature interaction: lift a new module onto the stack, feel the magnets seat it, close two latches, and watch the fabric absorb the capacity. No cables between modules, no configuration, no integrator. This section specifies how that experience is engineered without a single dishonest claim underneath it.
Three physics problems had to be resolved. First, retention: magnets alone cannot safely carry a 2.9 kg cantilevered module against a shear pull, so neodymium cone pairs handle alignment and self-seating within ±1.5 mm, and dual tool-free cam latches carry the structural load at a rated 25 kg shear (D/04). Second, signal integrity: the concept's pogo-pin interconnect cannot carry modern serial lanes reliably at over 10 GT/s across a separable interface, so the electrical join is a blind-mate high-speed connector of the ExaMAX class, recessed behind the magnetic registration so it can never be force-mated. Third, the pooling claim itself, resolved below.
What the customer experiences
- Seat it. The cones pull the module into registration; the connector blind-mates only once the faces are true.
- Latch it. Two cams close by hand. The stack is now one rigid, thermally aligned object.
- Watch it join. The fabric beacons, verifies the module's identity against its provenance certificate, and the console shows new capacity in under 90 seconds. Roles rebalance per the manifest: a second Pro becomes the dedicated deep-model runtime; a +S module becomes the replication target.
The pooling claim, made honest
The concept promised CXL.mem merged memory: a stack that behaves as one NUMA machine. Consumer-class silicon cannot do that today, and the claim would not survive one technical diligence call. The resolution (D/05–D/07): v1 delivers capacity pooling through workload placement over SiloLink-E, which is what buyers actually observe: more users, bigger models, faster ingestion. v2 activates SiloLink-X, a PCIe non-transparent bridge on the horizontal faces, enabling true tensor-parallel inference across adjacent modules, honoring the original design intent that horizontal means pooled compute, vertical means orchestration. CXL single-NUMA remains a v3 research gate, and is never marketed until it is real. The customer promise is stable across all three generations: click a module on, capacity goes up.
Software that treats the offline invariant
as a law of physics.
Every Silo Server runs the Silo Engine: an immutable, declaratively configured operating stack whose single governing rule is inherited from every product this company has shipped. Zero bytes egress by default. No telemetry, no accounts, no external APIs, no phone-home. The build system enforces it; the front lamp displays it.
The architecture extends the factory discipline already proven across the Silo app line: edit the manifest, re-emit everywhere. One file per machine, silo.server.json, declares the tier, the vertical template, the model pack, the role map, retention policy, and access control sources. From that one file the factory emits the NixOS configuration, the service topology, the compliance defaults, and the customer's provenance certificate. There is no hand-configured machine anywhere in the fleet, which is the only way a small team services a thousand appliances.
Model packs are content, not firmware
Models ship as signed, versioned packs, decoupled from the OS image. The default packs are built exclusively from permissively licensed weights, Apache-2.0 and MIT class, which removes redistribution ambiguity from the product entirely (D/12). Meta-licensed models remain available as customer-selected packs with license passthrough. Quarterly pack refreshes ride the SiloKey update path, so a sealed machine still gets a better brain four times a year.
| Lane | Class | Tier |
|---|---|---|
| FAST · interactive | 8B Q4 · permissive license | All |
| CORE · drafting · analysis | 32B-class Q4 | Studio · Pro |
| DEEP · reasoning · review | 70B-class Q4 · ≈42GB resident | Pro |
| EMBED | bge-large / nomic-class | All |
| VISION · OCR | 7B VLM assisting Tesseract | All |
Vertical templates
The tier is the hardware; the template is the firm. Selected at order time, a template overlays the engine with the industry's data model, retention policy, and prompt library:
- LEGAL · matter-centric corpus, privilege tagging, conflict-wall ACLs, citation-first drafting library, WORM off.
- RIA · client-account corpus, 17a-4 WORM retention on by default, books-and-records export.
- MEDICAL · patient-centric corpus, minimum-necessary ACL defaults, audit-log verbosity raised.
- ACCOUNTING · engagement-centric corpus, §7216 consent tracking, season-aware retention.
Ingestion is deliberately boring: drop files on the SMB share, SFTP, or an encrypted USB. Parse, OCR, chunk, embed, index. Documents keep their ACLs from the firm's directory (LDAP or AD), enforced as filters at query time, so an associate cannot retrieve across a conflict wall even by asking nicely.
Your records speak first.
The machine proves it.
Compliance marketing usually asserts conclusions. This product asserts mechanisms, and lets the buyer's counsel draw the conclusion, which is exactly how regulated buyers evaluate. Every claim below names the physical or cryptographic fact that produces it.
| Adversary · scenario | Cloud posture | Silo Server mechanism |
|---|---|---|
| Vendor breach · your provider is compromised | Your data is in the blast radius | There is no vendor holding data. Nothing to breach upstream. |
| Third-party legal process · subpoena served on a host | Provider may produce your data without you | All process must be served on the firm itself; counsel responds with full knowledge. |
| Physical theft · the box walks out | N/A | ZFS native encryption, key sealed in TPM 2.0 + admin passphrase; measured boot refuses tampered images; chassis fasteners tamper-evident. |
| Insider exfiltration | Provider-side logging, opaque | USB mass storage dead by default; only paired SiloKeys mount; every retrieval logged to an append-only dataset the admin cannot silently edit. |
| Supply-chain tamper · in transit | Opaque | Factory-attested image hash on the provenance certificate; first boot re-measures and must match before services start. |
| Vendor snooping · us | Terms-of-service trust | Zero-logical-access service model: Silo Systems holds no credentials to any deployed machine. Support runs on customer-exported, content-redacted diagnostic bundles. |
| Framework | The friction today | Silo Server posture |
|---|---|---|
| ABA Model Rule 1.6 · legal | Cloud AI use requires vendor diligence, disclosure analysis, sometimes client consent | Client data never transits third-party infrastructure; the confidentiality analysis collapses to the firm's own premises and personnel, the posture bar opinions treat most favorably. |
| HIPAA · medical | Every cloud AI vendor is a Business Associate; BAA chains sprawl | No cloud service exists in the data path, so no BAA chain exists for it. Where Silo performs on-site service, a BAA is offered; the default service model never touches PHI. |
| SEC 17a-4 · GLBA · finance | WORM storage and audit trails via cloud archive vendors | Retention-locked ZFS snapshot policy with immutability holds; audit logs on an append-only dataset. Designed to 17a-4(f); independent assessment scheduled Phase 3 (§X). Sold as designed-to until certified, in writing. |
| Data sovereignty · all | Foreign-process exposure via hosting providers | No hosting provider exists. Jurisdiction over the data is jurisdiction over the building. |
Support without custody: fleet homogeneity (one NixOS image) means most issues reproduce in the lab without ever seeing customer data. When a machine needs attention, the customer exports a diagnostic bundle to a SiloKey; the bundle contains system state and logs with document content structurally excluded, and its manifest lists every file it carries. A vendor that cannot access your data is a stronger compliance answer than a vendor that promises not to. This model is only possible because the fleet is immutable and manifest-driven; it is an architectural moat, not a policy.
Priced against the fear
it retires.
Pricing anchors to the buyer's alternative, not our BOM. A 20-seat firm's cloud AI bill runs $7,200 a year forever, with custody surrendered; an integrator-built private stack starts near $60,000. Silo Server prices between those poles, with hardware margin at appliance norms and recurring revenue attached to every unit (D/11).
| SKU | What it is | BOM est. | Price | Gross margin |
|---|---|---|---|---|
| Silo Server Mini | 1–3 seats · 32GB · ≤14B models | $1,510 | $3,499 | 57% |
| Silo Server Studio | 5–15 seats · 64GB · 32B-class | $1,960 | $5,499 | 64% |
| Silo Server Pro | 15–50 seats · 128GB · 70B resident | $2,890 | $8,999 | 68% |
| +C Compute module | Headless pooled worker · +64GB | $1,430 | $2,999 | 52% |
| +S Storage module | 8TB encrypted RAID-1 · replication target | $1,260 | $2,499 | 50% |
| +L Gateway module | Isolated SFP+ interfaces · multi-tenant LAN | $720 | $1,899 | 62% |
BOM basis: OEM Strix Halo mainboard, NVMe, extruded + CNC chassis and thermal assembly, pre-certified PSU module, SiloLink connector set, assembly and 48-hour burn-in labor, packaging with two SiloKeys. Margin improves up-tier because memory is the differentiator and memory is cheap relative to its price signal.
| Offer | Contents | Price |
|---|---|---|
| Silo Care | Quarterly signed update + model packs on SiloKey media · next-business-day support · hardware warranty maintained | $249 / mo · site |
| Silo Care+ | Care, plus 4-hour response window · advance-replacement hardware · annual on-site inspection · extended warranty to 5 years | $499 / mo · site |
| Silo Deploy | White-glove migration: corpus ingestion, directory integration, conflict-wall mapping, staff onboarding. The consulting practice, productized. | $2,500 – $7,500 · one-time |
The order experience
The configurator on the Silo Systems site is the product's front door, and it is deliberately shaped like buying a Mac:
- 01 · Size it. Seats and document volume recommend a tier. Override freely.
- 02 · Template it. Legal, RIA, Medical, Accounting. This writes the manifest.
- 03 · Extend it. +C, +S, +L modules; extra SiloKeys; Care tier; Deploy.
- 04 · Order it. 50% deposit funds procurement; balance on ship. Build-to-order, target 10 business days.
The deposit structure makes the business working-capital light: inventory is bought against orders, not forecasts.
What arrives in the crate
- The unit, sealed, image flashed and attested at the factory.
- Two paired SiloKeys, serialized to the chassis.
- The Certificate of Provenance: serials, manifest hash, image hash, burn-in record, signed. The firm's first compliance artifact exists before the machine is even powered on. A ledger copy is retained by Silo Systems.
- A one-page quick start. Plug into power and the firm switch, hold the front button, and the console appears on the LAN.
| Profile | Configuration | Initial | Recurring / yr |
|---|---|---|---|
| Solo practitioner | Mini + Care | $3,499 | $2,988 |
| 12-attorney firm | Studio + Deploy + Care | $8,999 | $2,988 |
| Lighthouse · 30-attorney firm | Pro + S + Deploy(4.5k) + Care+ | $15,998 | $5,988 |
| Growth event | Add Pro to existing stack | $8,999 | included |
A factory that fits
in a signature.
Every unit moves through one pipeline with four hard gates. Nothing ships around a gate. The pipeline is designed for a founder-led operation today and a contract manufacturer tomorrow, because every step is executed against the manifest and recorded to the ledger, the process itself is the training material for the first ops hire.
| Step | Responsible | Accountable | Consulted | Informed |
|---|---|---|---|---|
| Order intake · manifest | Configurator (automated) | Principal Architect | Customer | Assembly |
| Procurement | Principal Architect | Principal Architect | Board ODM · chassis shop | Customer (lead time) |
| Assembly · flash · burn-in | Contract assembly tech | Principal Architect | None | None |
| QA gate · G3 | Principal Architect | Principal Architect | None | Ledger |
| Ship · logistics | 3PL (milestone-gated) | Principal Architect | None | Customer |
| Activation · Deploy | Silo Deploy (services arm) | Principal Architect | Firm IT contact | None |
Stated plainly: v1 concentrates accountability in the founder by design, with two contract roles carrying labor. The first full-time operations hire is milestone-gated at a sustained 15 units per month, and this SOP is the job description.
| Scenario | Disposition |
|---|---|
| Board DOA at flash | Swap from spares bin (2 boards per 25 on order), RMA upstream; ledger notes serial substitution; lead time preserved. |
| Burn-in failure at G3 | Unit never ships. Tear-down analysis logged; component lot flagged; customer notified only if lead slips past D12. |
| Customer data migration exceeds Deploy scope | Change order under services rates; the appliance sale is never held hostage to the migration. |
| Field RMA | Advance replacement under Care+; drives are customer-retained on request (encrypted anyway); returned chassis wiped by key destruction, then re-imaged. |
| Lost SiloKey | Remaining paired key revokes the lost one; replacement key re-paired on premises. No remote revocation path exists, by design. |
Sell where the trust
already exists.
The beachhead is California law firms of 5–50 attorneys, reached first through the standing asset most hardware startups lack: seven years of consulting clients in exactly this segment.
Phase one · design partners. Three lighthouse deployments drawn from the services book, sold at a 40% pilot discount, never free: regulated buyers do not value free, and paid pilots produce reference-grade contracts. Deliverable: three named case studies with before/after workflow numbers and a compliance memo the buyer's counsel signed off.
Phase two · the configurator. Direct sales through the site, fed by the case studies, bar-association CLE talks, and the compliance-first content the consulting practice already produces. The founder's podcast and the legal-services network are distribution surfaces that already exist.
Phase three · channel. Legal-vertical MSPs and IT consultants, the people 20-attorney firms already trust, carry the product at 15–20% margin plus Deploy service revenue. The zero-access service model makes the channel comfortable: the box does not compete with their management contract.
| Quarter | Units | Hardware + services | Exit ARR (Care) |
|---|---|---|---|
| Q1 · pilots | 5 | $41k | $12k |
| Q2 · GA | 12 | $112k | $41k |
| Q3 | 22 | $204k | $92k |
| Q4 | 35 | $318k | $168k |
| Year 1 | 74 | $675k | $168k |
| Q5–Q8 | 260 | $2.4M | $700k |
Assumptions: blended initial order $8.6k rising to $9.7k with module attach; Care attach 80%, Care+ 25% of attach; Deploy attach 60% at $3.8k average; 3% quarterly churn on Care. Deposit-funded procurement holds inventory near zero. Gross margin blends to ≈61% hardware, ≈78% recurring.
| Method | Basis | Result |
|---|---|---|
| Bottoms-up · 3-year SOM | 0.5% of the 180k-site SAM · blended $12.4k initial + $3.4k/yr recurring | ≈ $11M initial · ≈ $3M ARR |
| Wedge capture · legal + RIA | 1% of ≈465k sites · $19k average site value at module attach | ≈ $88M hardware · ≈ $14M ARR |
| Ceiling | Every business that would rather own its infrastructure than rent it | The Apple model, aimed at the office |
Ninety days to a machine,
a year to a line.
| Window | Workstream | Gate · exit criterion |
|---|---|---|
| Days 1–30 | NixOS image sealed (egress audit clean), manifest → configuration.nix emitter, ingestion daemons, FastAPI spine on Strix Halo dev board | G-A · image boots sealed; 0 packets egress under 72h capture |
| Days 31–60 | Inference gateway + llama.cpp tuning (Vulkan/ROCm), full RAG pipeline, 10,000-page retrieval benchmark, Qdrant + Postgres + ZFS layout | G-B · 70B Q4 ≥ 5 tok/s sustained on Pro board; P95 retrieval < 900ms over 10k pages |
| Days 61–90 | Chassis first articles + thermal validation, SiloLink-E bring-up between two boards, K3s auto-join, factory flash pipeline end-to-end | G-C · two-module fabric joins < 90s; soak passes at spec acoustics |
| Months 4–6 | Pilot builds ×5, FCC/ETL testing, configurator live, Deploy runbooks, provenance ledger | G-D · 3 paying design partners activated at G4 |
| Months 7–9 | GA. First 25 production units, Care logistics (SiloKey mail cycle), channel pilot with one legal MSP | G-E · 25 units shipped · DOA < 2% · lead ≤ 10 BD held |
| Months 10–12 | SiloLink-X prototype (PCIe NTB tensor-parallel demo), 17a-4 assessment scoped, ops hire if ≥15 units/mo | G-F · 2-module tensor-parallel demo ≥ 1.6× single-node throughput |
| Risk | Severity | Counter |
|---|---|---|
| Silicon supply · Strix Halo allocation tightens | High | Dual-source across two ODMs; the manifest abstracts the board, so a board swap is a re-emit, not a redesign. Jetson Thor tracked as a priced fallback (D/01 revisit trigger). |
| SL-X signal integrity · PCIe over separable interface underdelivers | Med | v1 revenue does not depend on it; connector vendor eval boards before tooling; fallback is bonded dual-25GbE, which still doubles fabric bandwidth. |
| Founder bandwidth · solo founder, concurrent W-2 | High | Named head-on: full-time trigger is funding or 10 paid orders, whichever lands first; SOPs written so the first ops hire absorbs assembly and logistics; consulting book converts to Deploy revenue instead of competing for hours. |
| Support burden · appliances in the field | Med | One immutable image fleet-wide; zero-access diagnostics; A/B rollback means the worst field state is "previous known-good." Support cost modeled at $38/unit/yr inside Care COGS. |
| Compliance overreach · marketing writes a check legal can't cash | Med | Mechanism-stated claims only (§VI); 17a-4 sold as designed-to until independently assessed; template review by outside regulatory counsel before GA. |
| Model licensing · redistribution terms shift | Low | Default packs are Apache/MIT weights only (D/12); restricted-license models ship as customer-selected packs with passthrough. |
| Price resistance · $9k sticker at small firms | Low | Buyer math (TBL M-2) beats cloud by year two; leasing partner planned at GA so the lighthouse order lands under $650/mo, inside any firm's software budget. |
| Frontier-model envy · "the cloud model is smarter" | Med | Position honestly: this is the best model your data is allowed to touch. Quarterly packs compound quality; optional customer-enabled hybrid egress exists but is never the default and never silent (the lamp). |
Twelve forks,
closed in writing.
The record of the engineering and business judgment underneath this document. Each entry states the options weighed, the resolution, and the condition that would reopen it. Decisions without reopening triggers are dogma; these are not.
Options: Apple Silicon (best perf/watt, not licensable to OEMs); NVIDIA Jetson Thor (strong inference, CUDA/ARM lock-in, highest $/GB unified memory); AMD Strix Halo (128GB unified, x86, supply chain open to small integrators).
AMD Strix Halo across all tiers, differentiated by memory alone; integrate-don't-fabricate manufacturing, California assembly, CM gated at 500 units/yr.
Thor OEM module lands under $1,800 with 128GB-class memory, or Strix allocation fails two consecutive builds.
Pure passive caps sustained load near 25W and cannot serve a 70B model. The honest product is a specified acoustic contract, not a cooling ideology.
Vapor chamber + low-RPM assisted convection; ≤24 dBA idle, ≤32 dBA sustained at 1m; Mini passive-biased in typical duty.
A ≤40W SoC serves the Studio workload; Mini goes fully passive.
One module size across the entire line: 210 × 210 × 130 mm, 5.7L, 2.9 kg. Any face registers against any face; vent channels align across the stack. Half-modules rejected for v1: two SKUs of sheet metal is one too many.
Single grid module, all tiers and expansion modules identical externally.
A rack-adapter SKU is demanded by >10% of orders (a 19-inch tray holding two modules is the pre-designed answer).
Magnets cannot carry structural load; pogo pins cannot carry >10 GT/s across a separable interface reliably.
Neodymium cones align (±1.5mm, polarity-keyed); dual cam latches retain (25 kg shear); blind-mate ExaMAX-class connector carries signal, recessed so it mates only after registration.
Connector vendor qualification fails at PCIe G4 rates; fall back to bonded dual-25GbE on the X faces.
One protocol on every face: 25GbE switched, beacon discovery, K3s orchestration, roles from the manifest, auto-election by lowest serial. Orientation is ergonomics and thermals, not protocol, which makes the v1 product impossible to mis-assemble.
Uniform Ethernet-class fabric; join under 90 seconds; zero customer configuration.
Never; this remains the control plane even after SL-X ships.
Honors the founding intent: horizontal means pooled compute. PCIe Gen4 ×8 non-transparent bridging on the side faces enables tensor-parallel inference across adjacent modules.
SL-X specified into the v1 chassis (zone reserved, Studio/Pro populated) so v2 is a firmware-and-module story, not a new chassis.
G-F gate: ships only if the 2-module demo exceeds 1.6× single-node throughput.
The concept's merged-memory claim is not achievable on consumer-class silicon today and would not survive technical diligence.
Deferred to a v3 research gate, contingent on CXL-capable edge silicon. Never marketed until real. The customer promise ("click a module on, capacity goes up") is already true in v1 by workload placement.
CXL 3.x memory pooling appears in a ≤120W SoC roadmap from AMD or NVIDIA.
Options: NixOS vs custom Yocto. Yocto means maintaining a BSP, a team's worth of work. NixOS gives declarative single-source config (the manifest philosophy the whole company runs on), reproducible builds, and atomic A/B rollback.
NixOS immutable; silo.server.json → configuration.nix; quarterly signed updates on FIDO2-attested SiloKey media; optional customer-enabled network channel that turns the lamp white.
Fleet exceeds 2,000 units and OTA economics dominate; the signed-channel design already anticipates it.
vLLM's ROCm APU support is not mature enough to bet the fleet on; MLX is Apple-only and Apple is out (D/01).
llama.cpp (Vulkan/ROCm) behind an OpenAI-compatible gateway; RPC mode reserved for multi-node. The gateway isolates the engine so it is swappable without touching applications.
vLLM (or successor) demonstrates stable Strix Halo serving with ≥30% throughput gain on the 32B lane.
pgvector-only rejected: at multi-million-chunk corpora with per-document ACL filtering, a dedicated vector engine wins on filtered HNSW performance and operational isolation.
Qdrant (vectors) + PostgreSQL (application) + ZFS native encryption, hourly snapshots, replication by ZFS send to a +S module or an offsite Silo Server (Silo Mirror). Keys sealed in TPM 2.0 + admin passphrase.
Corpus profile at real customers stays under 500k chunks median for a year; revisit consolidation to shrink the surface.
Anchored to the buyer's alternatives (cloud per-seat forever vs $60k integrator build), not to BOM. Hardware carries appliance margin; the relationship carries recurring.
$3,499 / $5,499 / $8,999 base tiers at 57–68% gross margin; modules $1,899–$2,999; Care $249 and Care+ $499 monthly per site; Deploy $2,500–7,500. 50% deposit at order.
Pilot cohort closes at >80% without negotiation; raise Pro to $9,999 at GA.
Bundling Meta-licensed weights in a commercial appliance adds redistribution terms and attribution mechanics a regulated buyer's counsel will ask about.
Default packs are Apache-2.0/MIT weights exclusively (Qwen and Mistral class). Restricted-license models offered as customer-selected packs with license passthrough. The spec sheet's capability claims are written against the permissive default.
A restricted-license model becomes decisively superior for a vertical workload; ship it as a named optional pack, never the default.