Apple Intelligence Foundation Language Models
machinelearning.apple.com
machinelearning.apple.com
Performance in this first public release will be considered a baseline and can be improved over time. A PR disaster is a much bigger issue.
1. Two Main Models: - AFM-on-device: ~3 billion parameters, for efficient on-device use - AFM-server: Larger model for Private Cloud Compute
2. Architecture and Training: - Based on Transformer with optimizations - Three-stage training: core, continued, and context-lengthening - LoRA adapters for task-specific fine-tuning - Innovative quantization: 3.5-3.7 bits per weight
3. Performance and Benchmarks: - AFM-on-device outperforms larger models (e.g., Gemma-7B, Mistral-7B) - AFM-server competitive with GPT-3.5 - HELM MMLU (5-shot): AFM-on-device 61.4%, AFM-server 75.4% - GSM8K (8-shot CoT): AFM-server 83.3% - Strong in instruction-following (IFEval) - Best overall on Berkeley Function Calling Leaderboard
4. Capabilities: - Excels in instruction following, tool use, writing, math - Long context support up to 32k tokens - Specialized for tasks like summarization
5. Responsible AI: - Focus on user privacy and responsible AI principles - Extensive safety measures (red teaming, human evaluations) - Lower violation rates on safety prompts vs. other models
6. Unique Aspects: - "Accuracy-recovery adapters" post-quantization - Novel RLHF framework: "Iterative Teaching Committee" (iTeC) - New RL algorithm: MDLOO
It's more like recent those news, about boneless chicken can contain actually bones (and possibly can be made of non-chicken too).
Doing things on-device and not using users' data seems to be the most concrete thing.
If they * Show local performance * Provide a way for other LLM providers to receive compensation for providing local models
Then there are alot of use cases that get better with a local LLM instead of a remote LLM:
* Offline search
* Faster Text suggestion and transformation
* Client side AI assist for video games (This is better from a cost perspective for the game dev)
* Pushing the LLM runtime cost to the user