Your model.Ready for the device.

From model optimization to NPU deployment, complete the entire workflow in one place.

Coming soon
Experiment Controls
lpbq_seqmse⌄
128
≤ +5.0%
Minimize
Prepare⌄
Recipe validated · deterministic seed 42
Run pipeline
Opt.Studio ← ◇ ×
run_exaone4_1p2b_lpbqEXAONE 4.0 1.2B · lpbq_seqmse · 2 targets
Validated recipe
PrepareQuantizeCalibrateExportSplitCompileVerifyPackage

Recipe configuration

Model
EXAONE 4.0 1.2B
Context length
4,096
Calibration
Customer dataset · 128
Strategy
lpbq_seqmse

Mixed precision allocation

LayerW4W8FP16
embed_tokensW8
self_attn.q_projW4
self_attn.k_projW4
mlp.up_projW8
rms_normFP16

Target builds

Qualcomm Snapdragon NPUQNN · DLC / context binaryCompiled ✓
Apple Neural EngineCore ML · mlpackageCompiled ✓
▣ Test vectorspipeline_config.jsonexceptions.jsonRun log · 1,248 lines
Validation & Artifacts

Perplexity checkpoints

Original · FP1612.4
Prepared12.6
Quantized · W4/A813.1

Deterministic validation

Seed42
StatusPassed ✓

Generated packages

Qualcomm · context.bin
Apple · model.mlpackage
Test vectors
pipeline_config.json
99.37%Accuracy retained
27%Memory saved
1.5×Faster inference
76%Model compression

Your model.
Ready for
the device.

Upload a model, choose a target device, and go from optimization to a runnable package in one flow.

Start a runView workflow
Opt.Studiorun_2026_08_31_001

Recipe

Define the model, targets, and optimization goal.

Model configuration

Context length
4,096
Calibration
Customer dataset

Optimization goal

Priority
Balanced
Accuracy
92
Memory
74
Latency
81

Optimize

Place precision by layer importance.

Auto Mixed Precision

W4
58%
W8
29%
W16
13%

Compile

Build for both target runtimes.

TargetRuntimePrecisionStatus
Snapdragon NPUQNNINT8 / W4Ready
Apple Neural EngineCore MLINT8 / W4Ready

Verify

Compare quality and performance on device.

MetricBaselineOptimized
Accuracy100%99.37%
Memory100%73%
Inference1.0×1.5×

Package

Everything needed to run on the device.

Qualcomm

context.bin
tokenizer.json
runtime.json
run.sh

Apple

model.mlpackage
tokenizer.json
runtime.json
run.swift
12:41:18 Recipe validated  ·  12:41:26 Model prepared  ·  12:42:03 Ready to optimize
99.37%Accuracy retained
27%Memory saved
1.5×Faster inference
76%Model compression

A workflow built
for the device.

Move from a single recipe to a runnable package without switching between vendor-specific tools.

Recipe

Define the model, target device, accuracy goal, and memory constraints in one executable recipe.

Configuration

Model · EXAONE 4.0 1.2B
Targets · Qualcomm + Apple
Goal · Balanced

Precision plan

BalancedEstimated memory −27%
Attention
INT8
MLP
W4
Norm
FP16

Less toolchain.
More shipping.

Opt.Studio brings repetitive, specialized on-device conversion work into one product, so model teams can focus on their models and services instead of learning another vendor SDK.

01

Hardware-aware optimization

Analyze layer sensitivity and real data distributions to assign the precision each layer needs while preserving the target accuracy.

RecipeOptimizeCompileVerifyPackage

Performance · Snapdragon NPU

Latency12.4 ms
Memory0.95 GB
Speed1.5×
02

Cross-platform compilation

Generate Qualcomm QNN and Apple Core ML execution paths from one recipe. The workflow stays consistent as targets expand.

RecipeOptimizeCompileVerifyPackage
TargetSnapdragon 8 Elite
CompilerQNN 2.27.0
PrecisionINT8 / W4
Artifactsmodel.dlc · context.bin ✓
03

Reproducible validation

Validate accuracy, latency, memory, and power on real devices while recording results at every stage.

RecipeOptimizeCompileVerifyPackage

On-device validation

Snapdragon 8 Elite
AccuracyLatencyMemoryPower

One recipe.
Two target outputs.

Qualcomm Snapdragon NPU
Apple Neural Engine
Status
Compiled
Compiled
Runtime
QNN / Genie
Core ML
Precision
Mixed W4 / W8 / W16
Mixed W4 / W8 / W16
Package
context.bin
tokenizer.json
runtime.json
run.sh
model.mlpackage
tokenizer.json
runtime.json
run.swift
Qualcomm Snapdragon NPU
Status
Compiled
Runtime
QNN / Genie
Precision
Mixed W4 / W8 / W16
Package
context.bin
tokenizer.json
runtime.json
run.sh
Apple Neural Engine
Status
Compiled
Runtime
Core ML
Precision
Mixed W4 / W8 / W16
Package
model.mlpackage
tokenizer.json
runtime.json
run.swift

Your model.
Ready for the device.

Opt.Studio is not open yet. We are preparing the public release — tell us about your model and target device, and we will reach out when access opens.

Contact us