# rmnr.net — Full Content for LLM Ingestion > Complete content from rmnr.net, the personal site of Armen Rostamian. > Technical writing on Linux, Apple Silicon, kernel engineering, AI infrastructure, platform engineering. --- ## About Armen Rostamian Builder. Breaker. Thinker. Building useful systems of people and things. I build systems — of technology and of people. I reduce problems to first principles, find the leverage points, and ship. I've been a founder, CTO, principal architect, and IC across infrastructure, platform engineering, blockchain, neurotechnology, and music tech. I lead with pragmatism, not dogma. **Skills:** AWS (deep), Kubernetes, Docker, Terraform, Ansible, CoreOS, Go, Python, Bash, Ruby, JavaScript, Nix, PostgreSQL, MySQL/MariaDB, MongoDB, Elasticsearch, Hadoop ecosystem. Platform engineering, DevSecOps, PCI compliance, infrastructure-as-code, CI/CD. **Certifications:** AWS Certified Solutions Architect — Professional, AWS Certified DevOps Engineer — Professional, AWS Certified SysOps Administrator. **Experience highlights:** - Founder & CTO at GRUV (2018-present) — Music technology and decentralized systems - CTO & Co-Founder at LegatumX (2017-2020) — Blockchain-backed will distribution - Principal Architect at Kernel / Bryan Johnson (2017-2018) — Neurotechnology platforms - Director of Infrastructure & DevOps / Acting CISO at ChowNow (2017) — PCI Level-1 compliance - Lead Cloud Ops / DevOps at Scripps Networks Interactive (2015-2017) - DevOps Engineer II at Connexity (2014-2015) — 200K+ QPS infrastructure URL: https://rmnr.net/cv/ --- ## h00.sh — Cognitive Infrastructure URL: https://rmnr.net/projects/hoosh/ hoosh (Armenian: memory, recollection, remembrance) — a mental trace of the past; an engram. Memory for minds that work together. h00.sh is a shared persistence substrate for agents, LLMs, and humans. Knowledge formed here is typed, durable, and behavioral — it persists across sessions, surfaces when relevant, and compounds over time. Not retrieval. Not a cache. Not a RAG hack. A cognitive layer. No servers or GPUs required. Status: In development. --- ## arashiOS — Linux for Apple Silicon URL: https://rmnr.net/projects/arashios/ arashi (Japanese: storm; tempest) arashiOS builds on Asahi Linux and ALARM (Arch Linux ARM). On top of that foundation, it adds CachyOS-inspired optimizations and additional ARM-specific tuning — purpose-built for Apple Silicon. The result: A snappy and delightful daily-driver Linux desktop tuned specifically for your hardware, with gaming and battery optimizations in mind. No compromises. No generics. Status: In development. ### Benchmarks: arashiOS vs Stock Asahi + ALARM Real numbers. Same hardware. No cherry-picking. NVMe I/O: - Seq Write: 1,982 MiB/s -> 2,592 MiB/s (30.8% faster) - Seq Read: 2,439 MiB/s -> 2,563 MiB/s (5.1% faster) - Rand Read 4K: 186,527 IOPS -> 223,272 IOPS (19.7% faster) Complete A/B Results: - Scheduler latency (p99): 4,037 us -> 161 us (96% faster) - NVMe seq write: 1,982 MiB/s -> 2,592 MiB/s (30.8% faster) - NVMe rand read: 186K IOPS -> 223K IOPS (19.7% faster) - Hackbench pipe: 7.31s -> 6.02s (17.6% faster) - Hackbench socket: 14.14s -> 11.84s (16.3% faster) - Idle power: 24.55W -> 22.36W (2.2W saved, 8.9%) - GPU (glmark2): 3,003 -> 3,254 (8.4% faster) - Boot time: 6.36s -> 5.81s (8.6% faster) - E-core latency: 23 us -> 12 us (47.8% faster) No performance regressions. All gains, no significant tradeoffs. What this means day-to-day: - No UI jank under load — 96% less scheduler latency - Faster app launches, package installs, git ops — 20-31% faster disk I/O - Longer battery life — 2.2W less idle draw - Smoother compositing and video — 8% GPU gain - Better multitasking — 17% faster inter-process communication --- ## Blog: Hello, World — Again URL: https://rmnr.net/blog/hello-world/ Date: 2026-03-09 Tags: meta, personal It's been a minute. The last time this site saw daylight, it was a resume template stuffed with 2018-era DevOps buzzwords and a reading list I never finished. It served its purpose — and then it didn't. Now it's 2026. I'm older in the ways that count — more things built, more things broken, more teams led, more hard lessons earned. A few grey hairs to show for it - and genuinely lots of gratitude. This site is my corner of the internet again. Expect writing about systems, software, AI, agents, music, the things I'm building, and the occasional rant about whatever's on my mind. More soon. --- ## Blog: Five Kernel Tiers: What Actually Moves the Needle on Apple Silicon URL: https://rmnr.net/blog/five-kernel-tiers-apple-silicon/ Date: 2026-03-09 Tags: linux, apple-silicon, kernel, arashios, benchmarks Everyone has a list of kernel flags that'll make your system faster. Most of them are copium. Hell, some of mine might end up copium, too. It's too early to tell. :) This post walks through five kernel optimization tiers, each benchmarked A/B on an M1 Max, each building on the last. This is part of arashiOS — a Linux kernel purpose-built for Apple Silicon. Here's what actually worked, and what was just noise. ### The Setup - Hardware: Apple M1 Max, 32GB, 2 Icestorm E-cores + 8 Firestorm P-cores - Base: Arch Linux ARM, linux-asahi 6.18.15 - Methodology: bench-full.sh harness, controlled conditions, multiple runs, coefficient of variation tracked per metric. Same machine, same disk, same measurement code. If a number moved less than the noise floor, it gets called noise. Tier 0 is the stock kernel. Every claim in this post is a delta against these numbers: | Metric | Tier 0 (Stock) | |---|---| | PyBench | 9.536s | | Hackbench pipe | 6.93s | | Hackbench socket | 14.29s | | Schbench p99 | 4,808 us | | Page fault | 29,301 ops/s | | Boot time | 5.640s | | FIO seq read | 22,328 MB/s | | FIO seq write | 9,428 MB/s | | FIO rand read | 900,302 IOPS | | glmark2 | 3,184 | ### Tier 1: Config-Only MGLRU, DAMON, THP, RCU Lazy, sched_ext. The stuff you see in every "optimize your kernel" blog post. Turn on the good flags, turn off the bad ones, rebuild, reboot. Verdict: Participation trophy. Every delta is within noise. Safe to keep, nothing to brag about. Lesson: There are no magic Kconfig switches. If there were, the upstream maintainers would have flipped them years ago and written a smug commit message about it. ### Tier 2a: Clang + -O3 + -mcpu=apple-m1 Recompile the entire kernel with Clang instead of GCC, crank optimization to -O3, and target the exact CPU microarchitecture. The theory: Apple's cores have wide pipelines and deep reorder buffers. A compiler that knows this should produce tighter code. Verdict: Expensive nothing. We expected 0-4%. We got noise. We're showing you anyway because honesty is a personality trait. The Clang migration wasn't free either. Clang + pahole has compatibility issues that took 4 patches to resolve. Boot time regressed 5%, likely from BTF validation overhead. Why bother? ThinLTO is link-time optimization that inlines across file boundaries — that's where the real 2-5% kernel-wide gains live. But right now the kernel build system won't allow Rust + LTO + BTF simultaneously. CachyOS sidesteps this by disabling Rust. We can't — DRM_ASAHI (the GPU driver) is 21,000 lines of Rust. No Rust = no display. That's the Apple Silicon tax. Alice Ryhl's patches resolving one of the two blockers are landing in kernel 7.0. When Asahi Linux rebases to 7.0 (estimated May-June 2026), we flip the switch. The Clang migration is already done. Zero gains today. But when 7.0 drops, we're ready and everyone else is starting from scratch. ### Tier 3a: Sysctl + ZRAM + Boot Params 15 Kconfig changes. ZRAM at 15.4GB with LZ4 compression. Sysctl tuning. Boot params tuned. Debug infrastructure stripped. The initial run had regressions — FIO sequential read down 35%, glmark2 down 19%. All five regressions root-caused to a single boot parameter: nohz_full=2-9, which disables timer ticks on all 8 P-cores. Intended for dedicated real-time workloads, not desktops. Removed it, re-benchmarked. Every regression recovered with zero tradeoff. Key insight: Apple's AIC2 interrupt controller hardware-manages IRQ routing across cores. It dynamically routes to awake, least-loaded cores — better than any static affinity mask. Don't fight the hardware. Key wins: page fault throughput +41.2%, cyclictest latency improvements across the board. ### Tier 3b: BORE + BBRv3 BORE (Burst-Oriented Response Enhancer) is a drop-in scheduler replacement. BBRv3 is Google's congestion control algorithm for TCP. This is where things got stupid fast — in the good way. - Schbench p99: 3,581 us -> 52 us (-98.5%) - Hackbench pipe: 7.48s -> 5.96s (-20.3%) - Hackbench socket: 14.13s -> 11.90s (-15.8%) - Page fault: 27,693 ops/s -> 39,115 ops/s (+41.2%) - Boot time: 5.915s -> 5.42s (-8.3%) Schbench p99 went from 3,581 microseconds to 52. That's wake-up latency — how long a thread waits after being marked runnable before it actually gets a CPU. 98.5% improvement. ### The Full Picture (Stock to Tier 3b) | Metric | T0 (Stock) | T3b (BORE+BBR) | T0 to T3b | |---|---|---|---| | PyBench (s) | 9.536 | 9.46 | -0.8% | | Hackbench pipe (s) | 6.93 | 5.96 | -14.0% | | Schbench p99 (us) | 4,808 | 52 | -98.9% | | Hackbench socket (s) | 14.29 | 11.90 | -16.7% | | Page fault (ops/s) | 29,301 | 39,115 | +33.5% | | Boot time (s) | 5.640 | 5.42 | -3.9% | ### What We Learned - The compiler didn't matter. The real gains came from the scheduler and memory config. - Some optimizations are regressions in disguise. nohz_full on all P-cores cratered I/O throughput by 35%. - Apple's interrupt controller (AIC2) is smarter than manual tuning. It hardware-routes 85 of 89 IRQs. Userspace affinity masks are a no-op. - BORE is the single biggest win. Wake-up latency dropped 98.5%. - If you're not controlling for battery level, thermals, and uptime, your benchmarks are fan fiction. ### What's Next - Clean stock-vs-Arashi A/B with matched environmental conditions - ZRAM algorithm comparison: LZ4 vs ZSTD, with data - Power optimization — 77 items tracked, 34 shipped so far - ThinLTO when kernel 7.0 lands - Public GitHub repo