A CoreML conversion of Ultralytics' YOLOv8 nano model trained on Open Images V7, packaged for on-device inference on iPhone, iPad, and Apple Silicon Macs. Detects 601 object classes at 10 FPS average on iPhone 12 and...
Modellquelle
Quellenbeschreibung
A CoreML conversion of Ultralytics' YOLOv8 nano model trained on Open Images V7, packaged for on-device inference on iPhone, iPad, and Apple Silicon Macs. Detects 601 object classes at ~10 FPS average on iPhone 12 and newer.
This is the production model shipped inside — the same that runs in the App Store binary, byte-for-byte.
Quellen
1 QuelleVerifiziert 27. Sept.
Modellartefakte
2 Artefakteyolov8n_oiv7.mlpackage/Data/com.apple.CoreML/weights/weight.bin
bin · 6,72 MB · SHA-256 c648c2b9efc8…cefd · Hugging Face
HerunterladenQuellenauszüge
2 Auszüge.mlpackage| File | Format | Size | Purpose |
|---|---|---|---|
yolov8n_oiv7.mlpackage | CoreML (mlpackage) | 6.8 MB | The model — drop straight into an Xcode project |
yolov8n-oiv7.pt | PyTorch | 7.2 MB | Source weights from Ultralytics — for re-export to ONNX/TFLite/etc. |
class_names.txt | Plain text | 5.6 KB | All 601 OIV7 class labels, one per line, in model output order |
LICENSE | — | — | Dual GPL-3.0 / commercial license |
YOLOv8 + Open Images V7 gives you 601 detection classes — roughly 8× the 80-class COCO baseline that ships in most iOS object detection demos. Categories range from Accordion and Alpaca to Woodpecker, Wrench, and Zucchini. Far better fit for general-purpose camera apps than COCO.
Ultralytics distributes the original PyTorch weights (.pt), but no official CoreML build exists on the Hub for OIV7. This repo fills that gap so iOS developers can skip the conversion step and ship a 601-class detector in minutes — instead of the months it took to figure out the conversion the right way (see Conversion Notes below).
Measured inside RealTime AI Camera across iPhone X and newer:
| Device | Avg FPS | Notes |
|---|---|---|
| iPhone X | ~7-8 | A11 Bionic, no Neural Engine optimization for newer ops |
| iPhone 12 | ~10 | A14, Neural Engine accelerated |
| iPhone 14 Pro | ~12-15 | A16, fully optimized path |
| iPhone 15 Pro / 16 | ~15+ | A17 Pro / A18 |
Inference uses CoreML + Metal Performance Shaders + Apple Neural Engine on every supported chip. Frame rate auto-throttles based on thermal state to protect battery and avoid throttling spikes.
Drag yolov8n_oiv7.mlpackage into your Xcode project. Xcode auto-generates the yolov8n_oiv7 Swift class:
import CoreML
import Vision
let config = MLModelConfiguration()
config.computeUnits = .all // CPU + GPU + Neural Engine
let model = try yolov8n_oiv7(configuration: config)
let visionModel = try VNCoreMLModel(for: model.model)
let request = VNCoreMLRequest(model: visionModel) { request, error in
guard let results = request.results as? [VNRecognizedObjectObservation] else { return }
for obs in results {
print(obs.labels.first?.identifier ?? "?", obs.confidence, obs.boundingBox)
}
}
let handler = VNImageRequestHandler(cvPixelBuffer: pixelBuffer)
try handler.perform([request])
For a full production-grade integration with frame skipping, thermal management, object tracking, and Metal-accelerated preprocessing, see YOLOv8Processor.swift.
The .pt file is the standard Ultralytics format:
from ultralytics import YOLO
model = YOLO("yolov8n-oiv7.pt")
results = model("image.jpg")
results[0].show()
from ultralytics import YOLO
model = YOLO("yolov8n-oiv7.pt")
model.export(format="onnx") # ONNX
model.export(format="tflite") # TensorFlow Lite
model.export(format="coreml") # CoreML (will produce a similar .mlpackage)
601 categories from Open Images V7. See class_names.txt for the complete ordered list. A small sample:
Accordion, Aircraft, Airplane, Alpaca, Antelope, Backpack,
Banana, Bear, Bicycle, Boat, Book, Bottle, Bowl, Box, Bread,
... 580+ more ...
Woodpecker, Worm, Wrench, Zebra, Zucchini
The model output index matches the line number (0-indexed) in class_names.txt.
The naive path is
from ultralytics import YOLO; YOLO("yolov8n-oiv7.pt").export(format="coreml"). It produces a.mlpackagethat loads in Xcode, runs without errors, and is almost never what you actually want. This section documents what actually works in production after months of debugging.
If you want to skip all of this, just drop yolov8n_oiv7.mlpackage from this repo into your Xcode project. If you want to understand it (or convert your own variant — yolov8s, m, l, x), read on.
nms=True or you're writing the decoder yourselfProblem. A naive YOLOv8 export emits a raw tensor of shape [1, 605, 8400] for OIV7 (4 box coordinates + 601 class probabilities, across 8400 anchor positions). Vision's VNCoreMLRequest will not automatically give you VNRecognizedObjectObservations from this — it just hands you an MLMultiArray and you have to decode the boxes, run non-max suppression, and map class indices to labels in Swift. That's slow, error-prone, and burns CPU cycles you don't have at 30 FPS.
Fix. Export with NMS baked into the model so Vision returns VNRecognizedObjectObservations directly:
from ultralytics import YOLO
model = YOLO("yolov8n-oiv7.pt")
model.export(format="coreml", nms=True, imgsz=640)
nms=True is the difference between "I'm getting a 605×8400 multi-array, what do I do" and "Vision just gave me a clean array of detections with labels and confidences." Always use it for iOS.
Problem. "It exported, why doesn't it work" is almost always a tensor alignment bug. PyTorch and CoreML disagree about a lot of conventions, and the conversion tool silently picks defaults that often happen to be wrong. There are six independent things that all have to be right at the same time, and getting any one of them wrong gives you a model that loads, runs, and produces garbage — never an error message. You spend weeks thinking it's an accuracy problem when it's an alignment problem.
The six things, all of which must agree:
| # | Convention | PyTorch (YOLOv8) | CoreML / iOS Vision | Failure mode if wrong |
|---|---|---|---|---|
| a | Pixel value range | [0, 1] (normalized from uint8) | Whatever you set in ImageType(scale=...) | Zero detections everywhere — model sees pixel values 0–255 instead of 0–1 |
| b | Channel order | RGB | ImageType defaults to RGB but tracing can flip it | Detects color-insensitive classes only; fruit/lights/clothing wrong |
| c | Tensor layout | NCHW (batch, channels, H, W) | NHWC for ANE; NCHW for GPU/CPU fallback | Falls off the Neural Engine silently → 2-3× slower |
| d | Box coordinate origin | Top-left, normalized [0, 1] | Vision wants bottom-left normalized [0, 1] | Boxes appear upside-down in your overlay |
| e |
Fix. Specify everything explicitly at conversion time. Don't trust defaults. The full incantation that gets all six right for YOLOv8n on iOS:
import coremltools as ct
import torch
from ultralytics import YOLO
# Load PyTorch model and trace it
model = YOLO("yolov8n-oiv7.pt").model
model.eval()
example = torch.rand(1, 3, 640, 640) # NCHW, [0, 1]
traced = torch.jit.trace(model, example)
# Read class labels in model output order
labels = open("class_names.txt").read().strip().splitlines()
# Convert with EVERY alignment knob set explicitly
mlmodel = ct.convert(
traced,
inputs=[
ct.ImageType(
name="image",
shape=(1, 3, 640, 640),
scale=1/255.0, # (a) uint8 → [0, 1]
bias=[0, 0, 0], # YOLOv8 doesn't use ImageNet mean
color_layout=ct.colorlayout.RGB # (b) explicitly RGB, not BGR
)
],
classifier_config=ct.ClassifierConfig(class_labels=labels), # bake labels
compute_units=ct.ComputeUnit.ALL,
minimum_deployment_target=ct.target.iOS15,
convert_to="mlprogram", # mlpackage format, not legacy mlmodel
)
# Save with metadata
mlmodel.author = "Matt Macosko"
mlmodel.short_description = "YOLOv8n trained on Open Images V7 (601 classes)"
mlmodel.input_description["image"] = "Input image, will be resized to 640x640"
mlmodel.save("yolov8n_oiv7.mlpackage")
Then handle the box convention mismatch (d, e) in Swift after you get the Vision results — Vision actually flips the Y-axis for you when you use VNRecognizedObjectObservation, but only if the model was exported with nms=True (step 1). If you're decoding MLMultiArray output yourself, you have to flip Y manually:
// Vision gives you boundingBox with origin at bottom-left, normalized [0, 1].
// To draw on a UIView (top-left origin), flip Y:
let visionBox = observation.boundingBox // bottom-left origin
let viewBox = CGRect(
x: visionBox.minX * viewWidth,
y: (1 - visionBox.maxY) * viewHeight, // ← the Y flip
width: visionBox.width * viewWidth,
height: visionBox.height * viewHeight
)
This is the single biggest source of "I exported, the model loads, but everything is broken" — and there is no error. Just empty results, or boxes in the wrong place, or boxes for the wrong objects. Get all six right at once and the model snaps into working. Get any one wrong and you can't tell which one without methodically toggling each.
centerCrop is silently wrongProblem. YOLOv8 was trained on 640×640 inputs that were letterboxed (aspect-preserved with gray padding). Vision's default behavior on a VNImageRequestHandler is cropAndScaleOption = .centerCrop, which crops to the model's input size. On a 1920×1080 portrait camera frame, that means you lose the top and bottom of the image and detections at frame edges silently disappear. You won't get an error — you just notice your model is "kinda working but missing things."
Fix. Either:
cropAndScaleOption = .scaleFit on the request (CoreML will pad with black instead of cropping), ORIn RealTime AI Camera, MetalImageResizer.swift + Shader.metal do this — letterbox + BGRA→RGB + normalize in one Metal compute pass before the Vision request. Cuts preprocessing latency from ~3 ms to ~0.4 ms per frame.
Problem. AVCaptureVideoDataOutput defaults to kCVPixelFormatType_32BGRA. YOLOv8 was trained on RGB. If you feed BGRA directly into a model expecting RGB, you don't crash — you get a model that kind of works for color-insensitive classes (people, vehicles) and silently fails for color-sensitive ones (fruit, traffic lights, anything where red↔blue matters). You'll spend a...
--- license: gpl-3.0 tags: - coreml - yolov8 - yolov8n - object-detection - open-images-v7 - oiv7 - ios - iphone - apple-neural-engine - on-device - 601-classes library_name: coreml pipeline_tag: object-detection language: - en base_model: Ultralytics/YOLOv8 datasets: - google/open-images-v7 --- # YOLOv8n OIV7 — CoreML A **CoreML conversion** of [Ultralytics' YOLOv8 nano model trained on Open Images V7](https://docs.ultralytics.com/datasets/detect/open-images-v7/), packaged for on-device inference on **iPhone, iPad, and Apple Silicon Macs**. Detects **601 object classes** at ~10 FPS average on iPhone 12 and newer. This is the production model shipped inside [**RealTime AI Camera**](https://apps.apple.com/us/app/realtime-ai-cam/id6751230739) — the same `.mlpackage` that runs in the App Store binary, byte-for-byte. ## Files | File | Format | Size | Purpose | |---|---|---|---| | `yolov8n_oiv7.mlpackage` | **CoreML (mlpackage)** | 6.8 MB | The model — drop straight into an Xcode project | | `yolov8n-oiv7.pt` | PyTorch | 7.2 MB | Source weights from Ultralytics — for re-export to ONNX/TFLite/etc. | | `class_names.txt` | Plain text | 5.6 KB | All 601 OIV7 class labels, one per line, in model output order | | `LICENSE` | — | — | Dual GPL-3.0 / commercial license | ## Why this exists YOLOv8 + Open Images V7 gives you **601 detection classes** — roughly 8× the 80-class COCO baseline that ships in most iOS object detection demos. Categories range from `Accordion` and `Alpaca` to `Woodpecker`, `Wrench`, and `Zucchini`. Far better fit for general-purpose camera apps than COCO. Ultralytics distributes the original PyTorch weights (`.pt`), but **no official CoreML build exists on the Hub for OIV7**. This repo fills that gap so iOS developers can skip the conversion step and ship a 601-class detector in minutes — instead of the *months* it took to figure out the conversion the right way (see [Conversion Notes](#conversion-notes-pytorch--coreml-the-right-way) below). ## Performance Measured inside [RealTime AI Camera](https://github.com/nicedreamzapp/RealTimeAICam) across iPhone X and newer: | Device | Avg FPS | Notes | |---|---|---| | iPhone X | ~7-8 | A11 Bionic, no Neural Engine optimization for newer ops | | iPhone 12 | ~10 | A14, Neural Engine accelerated | | iPhone 14 Pro | ~12-15 | A16, fully optimized path | | iPhone 15 Pro / 16 | ~15+ | A17 Pro / A18...
Source context: 13 downloads · 1 likes · Pipeline object-detection · Library coreml · Repo divinetribe/yolov8n-oiv7-coreml
| Box format |
(cx, cy, w, h) center+size |
Vision wants (x, y, w, h) top-left+size |
| Boxes drift from objects, get larger than expected |
| f | Input type | torch.Tensor of shape [1, 3, 640, 640] | ImageType (not MultiArray) for Vision compatibility | Vision can't accept the model — you have to feed MLMultiArray manually |