Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/Splash engine the fastest local qwen38 on apple silicon
Dev48

© 2026 · All rights reserved.

Splash Engine - the fastest local Qwen3.8 on Apple Silicon

Источник: LM Studio Blog

Splash Engine - the fastest local Qwen3.8 on Apple Silicon

Source: LM Studio Blog

How to use Splash by Inco AI in LM Studio Bionic

September 26, 2026

What is Splash Engine?

Splash is an open-source inference engine from Inco AI for running language models locally on Apple silicon. It is optimized specifically for Qwen3.6-35B-A3B and Qwen3.8-27B. The engine provides GPU kernels and a memory plan tailored to each supported model. Each model ships with a dedicated DFlash 2 draft model for speculative decoding, which improves generation speed.

In Inco's tests on a 48 GB M5 Pro, Splash delivered roughly twice the decode speed of the next-fastest engine they measured on Qwen3.8-27B: 74 tokens per second on short prompts and 54 at 32K context. With four concurrent requests on short prompts, its combined throughput reached 170 tokens per second—3.9× the next-fastest engine in their comparison. Read more about Splash in Inco's blog post.

Use it in LM Studio Bionic

Download and install LM Studio Bionic 1.1.5 or newer, then open the app. Splash requires an M3-or-newer Mac running macOS 26.4 or later with at least 36 GB of unified memory; Inco recommends 48 GB or more.

Navigate to Settings > Runtime. Under Experimental backends, click Download next to Splash (Metal) to install the engine.

Download the Splash engine from Settings > Runtime.

Then go to Settings > Explore, paste one of the following Hugging Face links into the search bar, select the model, and click Download:

Once the download finishes, start a new session and select the model from the local model picker.

← All articles

More in Software Development

All →
Automattic has a new board after failed attempt to put CEO on leaveПресса
Automattic

Automattic has a new board after failed attempt to put CEO on leave

A new skill finds AI agent risks, fixes them, and proves the fix worked
Microsoft

A new skill finds AI agent risks, fixes them, and proves the fix worked

Some Supabase customers are publicly exposing reams of people’s data to the webПресса
Supabase

Some Supabase customers are publicly exposing reams of people’s data to the web

Blazor Basics: SEO Basics for Blazor Web Applications
Telerik

Blazor Basics: SEO Basics for Blazor Web Applications

Affected by layoffs? Don’t miss this $75 deal for your TechCrunch Disrupt 2026 Expo+ PassПресса
Expo

Affected by layoffs? Don’t miss this $75 deal for your TechCrunch Disrupt 2026 Expo+ Pass

Last 24 hours to save up to $200 on TechCrunch Disrupt 2026. Reason 5 of 5 to attend: MomentumПресса
Momentum

Last 24 hours to save up to $200 on TechCrunch Disrupt 2026. Reason 5 of 5 to attend: Momentum

More from LM Studio

Run Muse Glimmer locally
LM Studio

Run Muse Glimmer locally

Bionic now supports skills
LM Studio

Bionic now supports skills

How Auto Review works in Bionic
LM Studio

How Auto Review works in Bionic

Session References and Introspection
LM Studio

Session References and Introspection