Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/Splash engine the fastest local qwen38 on apple silicon 2
Dev48

© 2026 · All rights reserved.

Splash Engine - the fastest local Qwen3.8 on Apple Silicon

Источник: LM Studio Blog

Splash Engine - the fastest local Qwen3.8 on Apple Silicon

Source: LM Studio Blog

How to use Splash by Inco AI in LM Studio Bionic

September 29, 2026•Updated: September 29, 2026

What is Splash Engine?

Splash is an open-source inference engine from Inco AI for running language models locally on Apple silicon. It is optimized specifically for Qwen3.6-35B-A3B and Qwen3.8-27B. The engine provides GPU kernels and a memory plan tailored to each supported model. Each model ships with a dedicated DFlash 2 draft model for speculative decoding, which improves generation speed.

In Inco's tests on a 48 GB M5 Pro, Splash delivered roughly twice the decode speed of the next-fastest engine they measured on Qwen3.8-27B: 74 tokens per second on short prompts and 54 at 32K context. With four concurrent requests on short prompts, its combined throughput reached 170 tokens per second—3.9× the next-fastest engine in their comparison. Read more about Splash in Inco's blog post.

Use it in LM Studio Bionic

Download and install LM Studio Bionic 1.1.5 or newer, then open the app. Splash requires an M3-or-newer Mac running macOS 26.4 or later with at least 36 GB of unified memory; Inco recommends 48 GB or more.

Navigate to Settings > Runtime. Under Experimental backends, click Download next to Splash (Metal) to install the engine.

Download the Splash engine from Settings > Runtime.

Then go to Settings > Explore, paste one of the following Hugging Face links into the search bar, select the model, and click Download:

Once the download finishes, start a new session and select the model from the local model picker.

← All articles

More in Software Development

All →
Aurora CFO says 30,000 driverless trucks by 2030 isn’t as far-fetched as it soundsПресса
Aurora

Aurora CFO says 30,000 driverless trucks by 2030 isn’t as far-fetched as it sounds

Boeing 737 Max 10 certification delayed by software issue, FAA saysПресса
Boeing

Boeing 737 Max 10 certification delayed by software issue, FAA says

Shopify opens checkout to browser-based AI agents
Пресса
Shopify

Shopify opens checkout to browser-based AI agents

Air Teams: Bring Your Best Agentic Workflows to the Whole Team – and Automate Repeatable Work
JetBrains

Air Teams: Bring Your Best Agentic Workflows to the Whole Team – and Automate Repeatable Work

Rider 2026.2.3 Is Released!
JetBrains

Rider 2026.2.3 Is Released!

A More Reliable Compilation Scheme for Kotlin Multiplatform Modules
JetBrains

A More Reliable Compilation Scheme for Kotlin Multiplatform Modules

More from LM Studio

Run Muse Glimmer locally
LM Studio

Run Muse Glimmer locally

Bionic now supports skills
LM Studio

Bionic now supports skills

How Auto Review works in Bionic
LM Studio

How Auto Review works in Bionic

Session References and Introspection
LM Studio

Session References and Introspection