Local AIAI HardwareAI AgentsStartupsOpen-Weight Models

Ghost Core: What the $3,499 Personal AI Computer Can Actually Run

person
LaunchBoosts Research Desk·AI-assisted research
9 min read

Researched and drafted with AI assistance from the 7 public sources listed at the end of this article, then published after automated editorial checks. Spotted an error? Tell us at support@launchboosts.com.

Ghost Core: What the $3,499 Personal AI Computer Can Actually Run — LaunchBoosts

Ghost Core is a $3,499 desktop with no screen. It runs open-weight AI models and personal agents entirely on your own hardware, and you control it from your phone. Its San Francisco startup, Ghost, came out of stealth on October 5, 2026 with an $11 million seed round led by Andreessen Horowitz. The box is built around a 24 GB NVIDIA workstation GPU, which runs roughly 30-billion-parameter models well when they're quantized. It's a sensible privacy-first appliance. It isn't a replacement for frontier cloud models, and the price mostly pays for the software and integrations rather than the hardware.

Below: what Ghost actually announced, what the hardware can realistically run, what still goes over the internet, how Core compares with the alternatives, and what it means if you build agent software.

What Ghost announced

According to TechCrunch, Ghost was founded by Zain Javaid, a 19-year-old former quant trader, together with Nicholas Chua (19), Yifei Chen (24) and Gautam Sharda (23). Andreessen Horowitz led the $11 million seed, with Abstract, Audacious Ventures, SV Angel and Nova also investing. Javaid calls Core a "brain in a box": a dedicated machine that collects your data from your desktop, apps and smart home devices, and runs agents continuously so they can act as an assistant and carry out tasks for you.

Preorders opened on Monday, October 5. TechCrunch reported shipping in the last week of October, and Ghost's own product page gives the date as October 31, 2026 for the first batch, with free shipping. SiliconANGLE reported that the first batch sold out within hours. Ghost hasn't published how many units were in that batch.

The spec sheet

Component Ghost Core
GPU NVIDIA RTX PRO 4000 Blackwell SFF Edition, 24 GB GDDR7 ECC
GPU memory bandwidth 432 GB/s
AI compute (NVIDIA figure) 770 TOPS
CPU AMD Ryzen 5 7600 (6 cores / 12 threads)
System RAM 64 GB DDR5
Storage 1 TB NVMe SSD
Enclosure Brushed stainless steel
Interface Phone app, web app, voice; no display
Price $3,499, one-time
Support 1-year warranty, 30-day returns

Sources: Ghost, NVIDIA. NVIDIA lists the GPU at a 70 W maximum in a low-profile dual-slot form factor, which explains how Ghost fits it into a compact, quiet-looking box.

Preinstalled models

TechCrunch lists three models at launch: Qwen-3.8-Next, Qwen-3.8-27B and Gemma-4-31B. Ghost's product page adds a fourth, Muse-Glimmer-30B. You can also install models from Hugging Face or load models you trained yourself. Updates arrive over the air, and you can turn on automatic updates.

What 24 GB of GPU memory actually runs

Ghost hasn't published throughput numbers or the quantization levels it uses. The figures below are therefore back-of-envelope estimates from the published specs, not benchmarks.

A model's weights take up roughly (parameters × bits per weight ÷ 8) bytes. On top of that you need room for the KV cache, which grows with context length.

Model size 4-bit weights 8-bit weights Fits in 24 GB?
27B dense ~13.5 GB ~27 GB 4-bit yes, with room for context; 8-bit no
31B dense ~15.5 GB ~31 GB 4-bit yes; 8-bit no
70B dense ~35 GB ~70 GB Only with heavy CPU offload (slow)

Some practical consequences:

  1. The ~30B class is the sweet spot. That matches what Ghost ships. Expect quantized models, not full-precision ones.
  2. Long context eats memory. If an agent reads your whole inbox or a long document, the KV cache competes with the weights for the same 24 GB. Larger contexts mean either smaller models or more aggressive quantization.
  3. Bandwidth caps speed. For a dense model, generation can't produce tokens faster than memory bandwidth divided by the bytes read per token. At 432 GB/s with ~15.5 GB of 4-bit weights, the theoretical ceiling is roughly 28 tokens per second. Real-world speed will be lower. Mixture-of-experts models such as the "Next" variants only read their active parameters per token, so they can be much faster than their total size suggests. They do still need enough memory to hold all the weights, so some may spill into the 64 GB of system RAM.
  4. The always-on design suits background work. An agent that triages email overnight doesn't need chat-speed responses. Speed matters most for interactive voice use.

If you want a deeper look at running large open models on a single consumer GPU, our blog covers local inference trade-offs in more detail.

What stays local and what doesn't

Ghost's pitch is privacy and data ownership. According to TechCrunch, the company says:

  • All models and personal memory run on the device, and personal data is encrypted so that neither Ghost nor third parties can access it in the cloud.
  • You control how much autonomy agents have.
  • A firewall watches outgoing network requests to block unauthorized actions.
  • The device keeps working if Ghost shuts down. In the company's words, "All of the source code, logic, and model weights live on the device, not the services."

The caveats are just as important. Martin Cid Magazine notes that web search, email and over-the-air updates require an internet connection. So "no data leaves your home" can't be literally true for an agent that sends email or searches the web. What stays local is the inference and the memory store. The same report lists the integrations Ghost targets: email, calendars, finances, files, browser history, recordings, health records, smart home devices, and wearables such as Whoop, Oura and Eight Sleep. That makes the box a concentrated store of sensitive data. Its security depends on the firewall and on how agents are allowed to use credentials.

No independent security review of these claims has been published yet. Treat the firewall and encryption as vendor claims until researchers have tested shipping units.

Ghost Core vs. the alternatives

Core sells into a market where local AI hardware has become more expensive. On October 2, 2026, The Register reported that NVIDIA had launched a 64 GB DGX Spark at $4,999, and that the 128 GB model now lists at $6,950, nearly 75% more than a year earlier, because of a memory shortage. Both Spark models have 273 GB/s of memory bandwidth.

Option Upfront cost Model memory Bandwidth What you get
Ghost Core $3,499 24 GB VRAM + 64 GB system RAM 432 GB/s (GPU) Turnkey agent software, integrations, firewall, phone app
DIY PC with RTX PRO 4000 Card alone ~$3,100 new (early Oct 2026) Same Same Full control; you build the agent stack yourself
DGX Spark 64 GB $4,999 64 GB unified 273 GB/s Larger models fit, slower per-token reads; developer-oriented
DGX Spark 128 GB $6,950 128 GB unified 273 GB/s Runs 70B-class and larger; expensive
Cloud chatbot plan ~$20/month N/A N/A Frontier models; your data sits with the provider

The DIY price comes from GPUPrix's US price tracker, which listed a new RTX PRO 4000 Blackwell at $3,102 on October 3, 2026. That listing is for the standard card. The SFF Edition in Core is a separate SKU, and retail prices vary by seller.

How to read this:

  • The hardware is close to cost. With GPU prices where they are, the parts in Core add up to roughly what Ghost charges. Ghost isn't marking up silicon. You're paying for software and integration work.
  • The trade-off is capacity versus speed. Core has faster memory but much less of it than a DGX Spark. A Spark fits bigger models but reads each token more slowly. For ~30B models, Core's GPU should be faster. For anything above about 40B dense, Spark's capacity wins.
  • Cloud is still cheaper per unit of capability. Coverage points out that $3,499 equals about 14 years of a $20-a-month chatbot plan. A subscription also gets you frontier models that a 24 GB card can't run. The case for Core is privacy, ownership and always-on agents with access to your local data, not raw quality. To model your own API spend before deciding, try the LLM API Cost Calculator.

Questions to answer before you preorder

Ghost's FAQ and connectivity specs weren't fully visible on its product page when we checked. Get answers to these before you commit $3,499:

  1. Power and noise. It's meant to run 24/7. The GPU is capped at 70 W, but Ghost hasn't published whole-system power draw or noise figures.
  2. Quantization and context limits for each preinstalled model, and real tokens per second.
  3. Escalation to the cloud. Can you route hard tasks to a cloud model? If so, what data goes with the request?
  4. Credential handling. How do agents store OAuth tokens for email and banking, and what does the firewall actually block?
  5. Update policy. Automatic over-the-air updates change agent behavior on a machine that holds your financial and health data. Check whether you can pin versions.
  6. Availability. Coverage describes US pricing only, with no dates for Europe, Asia or Latin America.

What this means for builders

Even if you never buy one, Core shows where personal AI is heading.

  • A local-first agent runtime is now a product category. Hardware startups are betting that people will pay for an always-on, private agent. If you build agent tooling, check that it can run against a local OpenAI-compatible endpoint and a ~30B open model, not only frontier APIs.
  • Build for 24–32 GB of memory. Ghost shipping Qwen and Gemma models in the high-20B range suggests where on-device capability expectations are settling. Prompts, tool schemas and context budgets that work at that size will run on Core, on high-end gaming PCs and on workstations. The Token Counter helps you size context so it leaves room for the KV cache.
  • Outbound firewalls for agents are becoming a selling point. Ghost's design treats every outgoing request as something to inspect. That's a reasonable pattern for any agent with credentials, cloud-hosted or not. Default-deny egress, plus explicit user approval for actions that send data out, is worth copying.
  • Integrations are the moat, not the GPU. Ghost's value lives in connectors to email, calendars, finances and wearables. Startups building those connectors, or privacy-preserving memory layers, should note that a16z is funding this layer. Browse comparable products in our AI productivity tools directory.

What to watch next

  • Late October 2026: first units ship. Look for independent reviews that measure tokens per second, idle power draw and how agents behave on real inboxes.
  • Security research: the firewall and encryption claims need outside verification. A local box holding health, financial and email data is a valuable target.
  • Model updates: whether Ghost pushes newer open-weight releases over the air quickly, and whether larger mixture-of-experts models run well using the 64 GB of system RAM.
  • Memory prices: DGX Spark's price rises show that memory costs shape this whole category. If GPU prices stay high, Core's $3,499 price may be hard to hold for later batches. Ghost hasn't said anything about future pricing.

If you're evaluating now, don't preorder on the pitch alone. Wait for measured benchmarks, unless the 30-day return window is enough protection for you.

Frequently asked questions

How much does Ghost Core cost and when does it ship?

Ghost Core costs $3,499 one-time, with no subscription required according to launch coverage. Ghost's product page says the first batch ships October 31, 2026, with free shipping for that batch, a 30-day return window and a 1-year warranty. SiliconANGLE reported the first batch sold out within hours of preorders opening on October 5.

What hardware is inside Ghost Core?

Core uses an NVIDIA RTX PRO 4000 Blackwell SFF Edition GPU with 24 GB of GDDR7 ECC memory and 432 GB/s of bandwidth, an AMD Ryzen 5 7600 six-core CPU, 64 GB of DDR5 RAM and a 1 TB NVMe SSD. It has no screen and is controlled from a phone app, a web app, or by voice.

Does Ghost Core work without the internet?

Model inference and personal memory run on the device, so the core assistant works locally. Web search, email and over-the-air software updates still need an internet connection. Ghost also says the device keeps working if the company shuts down.

Which AI models can Ghost Core run?

It ships with open-weight models including Qwen 3.8-Next, Qwen 3.8-27B and Gemma 4-31B, and Ghost's product page also lists Muse-Glimmer-30B. You can add models from Hugging Face or load your own. With 24 GB of GPU memory, dense models around 30 billion parameters fit comfortably only when quantized to about 4–5 bits.

Is Ghost Core better than building my own local AI PC?

The hardware is not the selling point. The RTX PRO 4000 GPU alone has been listed at around $3,100 new in the US in early October 2026, so similar parts add up to roughly what Ghost charges. You're paying for the agent software, the integrations and the outbound firewall. If you only want to run local models, a DIY build or an existing workstation gives you more control.

Sources

  1. TechCrunch: At 19, founder raises $11M for Ghost, maker of a $3,499 computer for personal AI— techcrunch.com
  2. Ghost: Core product page and specifications— ghost.ai
  3. SiliconANGLE: AI agent hardware startup Ghost, led by its 19-year-old founder, raises $11M— siliconangle.com
  4. Martin Cid Magazine: Ghost's $3,499 Core has no screen, because the computer is for your AI agent— martincid.com
  5. NVIDIA: RTX PRO 4000 Blackwell SFF Edition specifications— nvidia.com
  6. The Register: Nvidia debuts $4,999 DGX Spark with half the RAM and storage amid memory crunch— theregister.com
  7. GPUPrix: RTX PRO 4000 Blackwell US price history— gpuprix.com
person

LaunchBoosts Research Desk

AI-assisted research

Explainers on software and AI industry trends, drafted with AI assistance from the public sources cited in each article and published after automated editorial checks for length, independent sourcing and originality.