Category: TECH NEWS MAIN

  • Apple shifts AI processing to the device

    Apple shifts AI processing to the device

    Apple handles most AI tasks on the device itself to keep user data private. The company uses its own silicon to run generative models locally on the iPhone, iPad, and Mac. This approach means the phone processes the information using its own processor without needing a cloud connection for basic tasks. For more complex requests that require larger foundation models, Apple uses Private Cloud Compute. This system uses custom-built server hardware in Apple data centers to run AI workloads. Apple designed this cloud system to ensure that personal user data remains inaccessible to anyone other than the user, including Apple staff.

    The A18 Pro chip powers this on-device intelligence. This 3-nanometer chip includes a 16-core neural engine with 35 TOPS of theoretical performance. It uses two Everest cores running at 4050 MHz and four Sawtooth cores running at 2420 MHz. These chips provide the 8 GB of RAM required to handle the intensive processing for Apple Intelligence. Earlier models like the iPhone 15 only have 6 GB of RAM, which might not be sufficient for these large language models.

    Component Specification
    Chipset A18 Pro
    Process 3 nanometers
    Neural Engine 16-core
    NPU Performance 35 TOPS
    RAM 8 GB LPDDR5X
    CPU Cores 6 (2 Everest, 4 Sawtooth)
    GPU Cores 6
    GPU Frequency 1490 MHz

    The MacBook Neo also uses the A18 Pro chip. This entry-level laptop costs $599 and includes a 16-core Neural Engine to support on-device AI tasks. It provides 16 hours of battery life and features a 1080p FaceTime HD camera. While the A18 Pro supports AI, the MacBook Neo lacks the M5 chips found in the MacBook Air. The M5 Pro and M5 Max chips handle AI tasks up to 4 times faster than the M4 predecessors.

    Siri AI adds new capabilities

    Siri AI acts as a more capable assistant through deep integration across Apple products. This new version of Siri draws on personal context to search messages, emails, and photos. It also performs systemwide app actions. Users can ask Siri to answer questions about content on their screen or search the web for up-to-date information. A dedicated Siri app allows users to restart or revisit conversations. iCloud syncs this conversational history privately across devices.

    The integration with third-party models changes how Siri functions. Apple uses Google’s Gemini models to power certain Apple Intelligence features. This collaboration helps Apple provide advanced chatbot capabilities. Apple is also discussing adding Google’s Gemini as an additional option later this year. This follows the existing partnership where ChatGPT assists with complex or creative requests. When Siri directs a question to ChatGPT, the user does not see the request history in their ChatGPT app.

    I find the lack of ChatGPT history frustrating.

    The writing tools in iOS 18 and macOS 15 provide immediate utility. These tools help users proofread, rewrite text in different tones, or summarize content. This functionality works well for professional or friendly messaging. However, some features like Genmoji take a long time to create. You might wait several minutes for a single emoji.

    Hardware requirements for intelligence

    Not every iPhone runs the full suite of Apple Intelligence features. The software requires specific hardware to manage the heavy processing loads. On-device AI tasks specifically require an iPhone 15 Pro or later. You also need an iPad with an M1 chip or later to run these features. The MacBook Neo, which uses the A18 Pro, also supports these capabilities.

    Device Model AI Compatibility
    iPhone 15 Pro Supported
    iPhone 15 Pro Max Supported
    iPhone 16 series Supported
    iPad (M1 or later) Supported
    MacBook Neo Supported
    iPhone 14 Not supported for on-device AI
    iPhone 13 Not supported for on-device AI

    The iPhone 16 series includes a camera control button. This button allows users to activate Visual Intelligence. This feature uses the camera and Siri to identify objects. This hardware-based approach differs from the iPhone SE, which is expected to use the A18 chip but will lack the camera control button. This means the SE will not have access to Visual Intelligence. The SE will instead feature a 6.1-inch OLED screen and Face ID, replacing the Home button.

    The A18 chip provides enough power for the upcoming iPhone SE. This device will likely cost $499. It includes USB-C and MagSafe for charging. While it lacks the camera control button, the A18 chip ensures the device can run Apple Intelligence.

    Privacy and cloud processing

    Apple relies on on-device processing to keep user data disaggregated. Data that stays on the device is not subject to centralized attacks. When the device needs more power, it sends requests to Private Cloud Compute. This system uses a hardened operating system to protect privacy. It uses code signing and sandboxing to limit the attack surface.

    Apple collaborates with Google and NVIDIA to expand this infrastructure. They use NVIDIA GPUs and Google Cloud to run more complex workloads. This expansion allows for agentic tool-use and complex reasoning. Apple maintains control over the software in these data centers. Devices only trust Private Cloud Compute software that Apple cryptographically approves.

    The company uses several methods to secure the cloud. It uses a verifiable, append-only ledger to track Google Cloud hardware in the PCC fleet. It also uses software attestation rooted in two separate independent vendors. This prevents attackers from easily accessing user data through the supply chain.

    Does the reliance on third-party data centers like Google Cloud weaken the privacy of Private Cloud Compute?

    The M5 Pro and M5 Max chips in the new MacBook Pro models also improve AI handling. These chips are designed for intensive tasks. The M5 Pro supports 64GB of unified memory. The M5 Max supports 128GB of unified memory. These higher memory capacities help the system manage the large models used in AI. The M5 Pro provides up to 30% better performance for pro workloads than the M4 Pro. The M5 Max provides up to 8x faster AI image generation than the M1 Pro.

  • Nvidia Blackwell Ultra supply bottleneck and manufacturing constraints

    Nvidia Blackwell Ultra supply bottleneck and manufacturing constraints

    Nvidia faces a manufacturing squeeze because TSMC’s CoWoS-L packaging capacity remains the primary constraint for Blackwell Ultra production. While TSMC aims to reach 130,000 CoWoS wafers per month by late 2026, the company remains fully booked with lead times between 52 and 78 weeks. Nvidia holds roughly 60% of this capacity, claiming 510,000 CoWoS wafers specifically for CoWoS-L. This concentration leaves Broadcom with 15% and AMD with 11%.

    The shortage stems from physical assembly needs rather than wafer fabrication. TSMC can etch more GPU dies than it can package. A GPU requires both a CoWoS slot and HBM stacks. Solving one alone moves nothing. Extra HBM with no packaging capacity results in inventory. Extra packaging capacity with no HBM results in idle line time.

    Thermal management issues also plague the Blackwell architecture. A mismatch in the coefficient of thermal expansion among the GPU chiplets, the LSI bridges, the RDL interposer, and the motherboard substrate causes warping and system failure. Nvidia had to redesign the top metal layers and bumps of the GPU silicon to improve yields. This redesign forces a requalification process with TSMC before mass production begins.

    Component 2026 Status Lead Time
    CoWoS-L Fully booked 52 – 78 weeks
    HBM3e Sold out N/A
    HBM4 Ramping N/A
    N3 Logic Tight 52 – 78 weeks

    Memory and hyperscaler demand

    High Bandwidth Memory (HBM) availability creates a second bottleneck. HBM3e is sold out for 2026, with prices increasing by double digits year-over-year. SK Hynix supplies approximately 62% of Nvidia’s HBM4 and roughly two-thirds of its HBM3e. This supply concentration forces Nvidia to rely on a single primary partner to meet its roadmap.

    Microsoft is attempting to reduce its dependence on Nvidia by building custom silicon. Microsoft is in discussions with TSMC to secure manufacturing capacity for over 300,000 Maia 300 chips for 2027 delivery. This follows the January launch of the Maia 200. Microsoft aims to produce gigawatts of capacity through these custom chips.

    Large cloud providers also consume the remaining supply through massive forward orders. Microsoft, Google, Meta, and Amazon placed multi-billion-dollar orders for Blackwell GPUs in 2025. These orders consume most of the available allocation through 2026 and 2027. These commitments crowd out mid-market and enterprise customers who previously bought through standard channels.

    I find the reliance on a single packaging technology for the entire roadmap risky. If TSMC cannot resolve the warping issues in CoWoS-L, the entire Blackwell ramp stalls.

    Blackwell Ultra specifications

    Blackwell Ultra targets the AI factory market using a dual-reticle design. This design connects two reticle-sized dies using the NVIDIA High-Bandwidth Interface, which provides 10 TB/s of bandwidth. The chip utilizes TSMC 4NP manufacturing and contains 208 billion transistors.

    The memory subsystem in Blackwell Ultra provides 288 GB of HBM3e per GPU. This capacity represents a 3.6x increase over the H100 and a 50% increase over the original Blackwell. The total bandwidth reaches 8 TB/s per GPU, which is a 2.4x improvement over the H100’s 3.35 TB/s.

    Feature Blackwell Ultra Spec
    Transistor Count 208B
    Memory Capacity 288 GB HBM3e
    Memory Bandwidth 8 TB/s
    Tensor Cores 640 (5th Gen)
    NVFP4 Compute 15 PetaFLOPS
    TDP 1,400W

    The architecture includes 160 Streaming Multiprocessors organized into eight Graphics Processing Clusters. Each SM contains four fifth-generation Tensor Cores. These cores use the second-generation Transformer Engine to handle NVFP4 precision. This format reduces the memory footprint by 1.8x compared to FP8.

    The Blackwell Ultra also doubles the throughput for key instructions in the attention layer. This change allows for 2x faster attention-layer compute compared to previous Blackwell models. This modification targets reasoning models that use large context windows.

    The competitive landscape

    AMD and custom silicon designs challenge Nvidia’s market position. AMD’s MI325X delivers competitive results against Nvidia’s H100 for certain inference workloads. AMD’s chips deliver 40% more tokens per dollar on LLM inference workloads compared to Nvidia.

    Hyperscalers are moving toward self-sufficiency to lower costs. Microsoft’s Maia 200 offers 30% better performance per dollar than the latest generation hardware in its current fleet. Google uses its own TPUs, and Amazon utilizes its own Trainium and Inferentia chips. These custom solutions reduce the high-margin sales available to Nvidia from its largest customers.

    The software ecosystem remains Nvidia’s main defense. CUDA provides the programming model for GPU-accelerated computing and has a two-decade head start. Every major deep learning framework like PyTorch and TensorFlow uses CUDA as a native backend. Switching to AMD’s ROCm requires migrating every application, library, and operational workflow.

    Can Nvidia maintain its valuation if custom silicon replaces its primary customers?

    The market currently views the Blackwell supply issues as a temporary manufacturing hurdle. However, the combination of TSMC packaging limits and the aggressive custom silicon programs at Microsoft and Google creates a difficult environment for Nvidia to sustain its growth rates. The company must manage the transition to CoWoS-L without letting its leading-edge customers migrate to in-house alternatives.