INCOMING PHD · TENNESSEE TECH · FALL 2026 · BAY AREA, CA

Amr Akmal Abouelmagd

M.S. Computer Science / HPC & GPU Systems Researcher / Software Engineer

My research runs from GPU architecture, partitioning, and power on the AMD MI300A APU behind El Capitan — the world's fastest supercomputer — to fault tolerance in HPC, KV-caching architecture and performance, and reshaping MPI for the AI era so it stays fault-tolerant and keeps making strong progress under failure. This fall I begin my PhD journey with Prof. Anthony Skjellum.

§ 01

News

// recent
Jan 2026
Started a continued research internship at LLNL's Computation Directorate, extending GPU-architecture work on the AMD MI300A.
Dec 2025
Paper “GPU Partitioning, Power, and Performance of the AMD MI300A” accepted at SCA / HPC Asia 2026.
Nov 2025
Top 5 Finalist in the SC'25 Graduate Student Research Competition for MI300A APU performance analysis.
Aug 2025
Wrapped up a summer research internship at LLNL benchmarking RAJA-Suite kernels across MI300A partitioning modes.
2024
Received the Outstanding Student Paper Award at IEEE HPEC 2024.
§ 02

Research interests

// focus areas
01GPU Architecture, Partitioning & Power
02Fault Tolerance in HPC Systems
03KV-Caching Architecture & Performance
04MPI for AI — Fault-Tolerant, Strong Progress
05Distributed Systems & Parallel Computing
06Systems Performance Analysis
§ 03

Publications

// 7 papers · 4 first-author
01
GPU Partitioning, Power, and Performance of the AMD MI300A
A. Abouelmagd, D. Boehme, S. Brink, J. Burmark, M. McKinsey, A. Skjellum, O. Pearce
SCA / HPC Asia 2026 First author
02
Using Hardware Metrics to Understand Performance of RAJA Suite Kernels in Different GPU Modes on MI300A
A. A. Abouelmagd, O. Pearce, S. Brink, M. McKinsey, D. Boehme, J. Burmark, B. Ryujin, T. Scogland, A. Skjellum
Poster · SC'25 First author Top 5 Finalist · GSRC
03
Load Imbalance in HPC Applications: Improved Profiling and New Ways to Use Wasted Cycles
S. Yang, X. Yao, G. Nansamba, A. A. Abouelmagd, A. Skjellum, M. Herbordt
IEEE HPEC 2025
04
A Survey of Optimization Approaches for MPI Alltoall and MPI Alltoallv
E. Namugwanya, G. Nansamba, A. A. Abouelmagd, A. Skjellum
SAI Computing Conference 2025
05
Leveraging the Power of AI and Social Interactions to Restore Trust in Public Polls
A. A. Abouelmagd, A. Hilal
CSCI 2025 First author
06
Emerging Paradigms for Securing Federated Learning Systems
A. A. Abouelmagd, A. Hilal
IEEE GCAIoT 2025 First author
07
Cycle-Stealing in Load-Imbalanced HPC Applications
P. H. Chen, A. Bali, S. Yang, P. Haghi, C. Knox, B. Li, A. A. Abouelmagd, A. Skjellum, M. Herbordt
IEEE HPEC 2024 Outstanding Student Paper
§ 04

Experience

// research + industry
Lawrence Livermore National Lab
Computation Directorate · Livermore, CA
Jan 2026 — Aug 2026
Research Intern
  • Continuing GPU-architecture research on the AMD MI300A, applying performance profiling and runtime analysis to optimize HPC application throughput and resource utilization across large-scale parallel workloads.
Lawrence Livermore National Lab
Computation Directorate · Livermore, CA
May 2025 — Aug 2025
Research Intern
  • Authored a peer-reviewed paper on MI300A partitioning, power, and performance (accepted at SCA/HPC Asia 2026), targeting optimization of El Capitan, the world's fastest supercomputer.
  • Earned a Top 5 Finalist placement in the SC'25 Graduate Student Research Competition.
  • Conducted systems-level research on the MI300A APU — GPU runtime scheduling, dynamic power sharing, and heterogeneous memory behavior.
  • Designed benchmarking experiments across SPX, TPX, and CPX partitioning modes using RAJA Performance Suite kernels, surfacing up to 30% execution-time variance.
Incorta
Alexandria, Egypt
Aug 2021 — Feb 2024
Software Engineer II → R&D Engineer → Graduate Intern
  • Optimized a distributed analytics platform on Kubernetes + ZooKeeper, improving scalability and query throughput for enterprise workloads.
  • Delivered a 2× indexing speedup and 7× query-latency reduction on datasets exceeding 1 billion records, improving SLA compliance.
  • Designed an internal caching layer that cut cloud object-storage I/O and infrastructure cost with no hardware changes.
  • Resolved CPU bottlenecks, raising average utilization from 40% to 85% and reducing over-provisioning.
  • Improved monitoring and telemetry, reducing incident MTTR across microservices.
JavaPythonGoKubernetesZooKeeperMySQLDistributed Systems
§ 05

Education

// incoming PhD
Tennessee Technological University
Cookeville, TN
Fall 2026 — Incoming
Ph.D. in Computer Science
  • Incoming doctoral researcher advised by Prof. Anthony Skjellum, focusing on high-performance computing and GPU systems.
Tennessee Technological University
Cookeville, TN
Jan 2024 — Summer 2026
M.S. in Computer Science
  • GPA 4.0 / 4.0. Research in high-performance computing, GPU performance analysis, and distributed systems.
Alexandria University
Faculty of Engineering · Alexandria, Egypt
Sep 2016 — Jun 2021
B.S. in Computer Engineering
  • Foundations in computer architecture, systems, and parallel programming.
§ 06

Selected projects

// systems / parallel

BusTub Database Engine

Buffer-pool manager with page-replacement policies, plus a thread-safe LRU-K eviction algorithm that maximizes memory utilization under concurrent workloads.

C++MakefileCMU

CUDA Parallel Softmax

Optimized softmax kernel using shared memory and block-level reduction to cut global-memory latency; throughput scaled via memory coalescing and parallel row-wise processing.

C++CUDAGPU

Distributed Key-Value Store

Leader-coordinated store with auto-partitioning and live rebalancing. Fault tolerance via the Bully election algorithm for zero-data-loss failover; multithreaded reads with mutexes and atomics.

C++MultithreadingDistributed
§ 07

Skills & tooling

Languages

C / C++PythonJavaGo

Parallel & HPC

MPICUDAKokkosRAJA Perf Suite

Systems & Infra

KubernetesDockerZooKeeperMySQL

Developer tools

LinuxGitCMakeMakefile
§ 08

Awards & achievements

Top 5 · Graduate Student Research CompetitionACM/IEEE SuperComputing (SC'25)
2025
Outstanding Student Paper AwardIEEE HPEC 2024
2024
10th Place · AlexCPC — qualified for ECPCEgyptian Collegiate Programming Contest
2018
Cloud DevOps NanodegreeUdacity