Research library
Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.
Search and filter
Paper list
T13. Edge AI, streaming inference, and resource scheduling
Running models near the machine under limited compute and latency budgets.
- T13-9Candidate approach
INFaaS
- T13-10Candidate approach
Orca: A Distributed Serving System for Transformer-Based Generative Models
- T13-11Candidate approach
Efficient Memory Management for Large Language Model Serving with PagedAttention
- T13-12Candidate approach
Gandiva: Introspective Cluster Scheduling for Deep Learning
- T13-13Candidate approach
Tiresias: A GPU Cluster Manager for Distributed Deep Learning
- T13-14Candidate approach
AntMan: Dynamic Scaling on GPU Clusters for Deep Learning
- T13-15Candidate approach
Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads
- T13-16Candidate approach
Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
From research to product use
Operating capabilities, pilots, and technologies in development are identified separately.
View technology