Personal AI should feel personal: responsive, quiet, and close. Puppy One pairs the new generation of Apple silicon with 🤫 Private Agent One - guided day-one activation included, consent-led, and tuned for the work you actually do.
The strongest balance of model headroom, bandwidth, power, and capital cost for demanding personal AI.
04Local frontier
Configured on Mac Studio · M5 Ultra
Puppy One 256 / 512
Very large open-weight models, research, multi-agent teams, and clusters.
Our build
36-core CPU · 80-core GPU · 256GB or 512GB · 2TB+
Memory
Up to 512GB
Bandwidth
1.2TB/s
Buy it for memory fit and absolute local capability - not because the biggest chip automatically means the best efficiency.
A better efficiency standard
The fastest token is not always the most useful one.
Agentic systems read, reason, retrieve, call tools, wait, and act. A single tokens-per-watt number hides that journey. We will publish the number - and the outcome behind it.
01
Tokens per joule
Prefill and decode measured separately at the wall, not inferred from chip marketing.
02
Tasks per kilowatt-hour
Successful, human-accepted agent jobs per unit of energy - our primary outcome metric.
03
Local completion rate
The share of useful work completed without sending private context away from the device.
04
Time to useful
Prompt processing, first token, final answer, and tool latency measured as one experience.
🤫 research note · August 2026
The smallest machine that completes the job wins.
Apple's new systems change the local-AI frontier in two different ways. M6 is the efficiency story: a 2-nanometer design, dual Neural Engines, GPU Neural Accelerators, and up to 170GB/s in a five-inch desktop. M5 Ultra is the capacity story: up to 512GB of unified memory and 1.2TB/s of bandwidth, enough to keep models with hundreds of billions of parameters on one desk.
Between them, M5 Pro and M5 Max are where most professional customers should live. The 64GB mini is the compact 32B machine. The 128GB Studio is our high-confidence recommendation for 70B-class models and concurrent agents. Ultra earns its place only when the workload needs its memory.
Apple reports relative "LLM prompt processing" gains, but its product pages define that test as time to first token. Apple does not publish the tested model, quantization, context length, sustained decode rate, or workload wall power. Its technical specifications publish maximum continuous system power - 155W for Mac mini and 480W for Mac Studio - not inference power. Any exact Puppy token-per-watt figure today would be false precision.
Our launch qualification will record wall energy, time-to-first-token, prompt and decode throughput, thermals, quality, tool success, and idle draw for every recommended model/configuration pair. Until that matrix passes, efficiency labels on this page are recommendations, not benchmark claims.
Reserve any Puppy for any amount from $0.01. Your reservation is refundable and credited toward purchase.
And a straight deal on price: before you pay, hussh checks the lowest advertised U.S. price for your exact Apple configuration - Apple, Costco, and the major retailers - and prices your hardware below it.2