Awesome System Papers Wiki
Search
搜索
暗色模式
亮色模式
探索
标签: area/ai-infra
此标签下有64条笔记。
2026年8月20日
APE-ICLR25
context-augmented-generation
kv-cache
rag
parallel-encoding
long-context
area/ai-infra
2026年8月20日
AVO-arXiv26
agentic-search
gpu-kernels
evolutionary-search
attention
blackwell
area/ai-infra
domain/auto-research
2026年8月20日
AdaExplore-arXiv26
gpu-kernels
coding-agent
test-time-adaptation
search
triton
area/ai-infra
domain/auto-research
2026年8月20日
Agentix-NSDI26
llm-agents
agent-serving
program-scheduling
preemption
long-horizon
area/ai-infra
area/agent-systems
2026年8月20日
AttnRes-arXiv26
llm-architecture
residual-connections
attention
ml-systems
inference
area/ai-infra
2026年8月20日
Axe-arXiv26
ml-compiler
tensor-layout
gpu
distributed-computing
dsl
area/ai-infra
2026年8月20日
BlendServe-ASPLOS26
llm-inference
offline-serving
batching
prefix-caching
resource-overlap
area/ai-infra
2026年8月20日
BlitzScale-OSDI25
llm-serving
autoscaling
model-as-a-service
multicast
serverless
area/ai-infra
2026年8月20日
CAKE-arXiv26
compiler-agent-codesign
gpu-kernels
kernel-generation
ir
blackwell
area/ai-infra
domain/auto-research
2026年8月20日
CacheBlend-EuroSys25
llm-serving
rag
kv-cache
cache-reuse
selective-recompute
prefix-caching
area/ai-infra
2026年8月20日
CacheGen-SIGCOMM24
llm-serving
kv-cache
compression
streaming
long-context
network
area/ai-infra
2026年8月20日
Cartridges-ICLR26
llm-inference
kv-cache
long-context
context-distillation
prefix-tuning
synthetic-data
area/ai-infra
2026年8月20日
CoX-MoE-DAC26
llm-inference
moe
cpu-gpu
amx
expert-offloading
throughput
area/ai-infra
2026年8月20日
ContextAwareMoE-CXLNDP-arXiv25
llm-inference
moe
cxl
ndp
quantization
expert-offloading
area/ai-infra
2026年8月20日
DiffKV-SOSP25
kv-cache
llm-serving
compression
gpu-memory
quantization
area/ai-infra
2026年8月20日
EventTensor-MLSys26
compiler
megakernel
llm-inference
moe
gpu-scheduling
area/ai-infra
2026年8月20日
FlashInfer-Bench-MLSys26
gpu-kernels
llm-inference
benchmark
agent
flashinfer
area/ai-infra
domain/auto-research
2026年8月20日
FlashInfer-MLSys25
llm-inference
attention
gpu-kernels
kv-cache
jit
area/ai-infra
2026年8月20日
FlowANN-OSDI26
vector-search
anns
gpu
graph
cpu-gpu-offloading
area/ai-infra
2026年8月20日
FluxMoE-arXiv26
moe
llm-inference
kv-cache
expert-offloading
memory-management
lossless-compression
area/ai-infra
2026年8月20日
GraphPipe-ASPLOS25
distributed-training
pipeline-parallelism
dag
scheduling
model-parallelism
area/ai-infra
2026年8月20日
He-GPUKernelFusion-SOSP26
gpu
kernel-fusion
dynamic-workload
sm-cooperation
area/ai-infra
2026年8月20日
HeteroInfer-SOSP25
mobile-llm
npu
gpu
heterogeneous-computing
soc
area/ai-infra
2026年8月20日
IceCache-arXiv26
llm-inference
kv-cache
long-context
offloading
sparse-attention
memory-management
area/ai-infra
2026年8月20日
KVCacheInTheWild-ATC25
llm-serving
kv-cache
prefix-caching
workload-characterization
cache-eviction
production-traces
area/ai-infra
CPU
2026年8月20日
LLMQueryReordering-MLSys25
llm-inference
data-analytics
prefix-caching
query-optimization
relational-data
area/ai-infra
2026年8月20日
LLMSteer-NeurIPSW24
llm-inference
kv-cache
prefix-caching
attention-steering
long-context
area/ai-infra
2026年8月20日
LMCache-arXiv25
llm-inference
kv-cache
prefix-caching
disaggregation
cache-layer
production-systems
area/ai-infra
1/2
2026年8月20日
LMetric-OSDI26
llm-serving
request-scheduling
kv-cache
load-balancing
area/ai-infra
2026年8月20日
LatencyOptimal-MoELB-INET4AI25
moe
llm-inference
expert-parallelism
load-balancing
ilp
gpu
area/ai-infra
2026年8月20日
Libra-ICLR26
moe
llm-inference
load-balancing
expert-parallelism
prefill
area/ai-infra
2026年8月20日
MOE-INFINITY-arXiv24
llm-inference
moe
expert-cache
offloading
personal-computing
area/ai-infra
2026年8月20日
MPK-OSDI26
gpu
compiler
mega-kernel
llm-inference
tensor-program
area/ai-infra
2026年8月20日
MSA-arXiv26
llm-inference
long-context
sparse-attention
kv-cache
memory-systems
rag
area/ai-infra
2026年8月20日
MagicDec-ICLR25
speculative-decoding
long-context
kv-cache
llm-serving
inference
area/ai-infra
2026年8月20日
Miao-LLMServingSurvey-CSUR26
survey
llm-serving
inference
systems
algorithms
area/ai-infra
2026年8月20日
MoE-Lightning-ASPLOS25
moe
llm-inference
cpu-gpu-pipeline
offloading
performance-model
area/ai-infra
2026年8月20日
MoE-nD-arXiv26
llm-inference
kv-cache
compression
quantization
long-context
routing
area/ai-infra
2026年8月20日
Multiverse-NeurIPS25
parallel-generation
reasoning
non-autoregressive
llm-inference
model-system-codesign
area/ai-infra
2026年8月20日
NEO-MLSys25
llm-inference
cpu-offloading
kv-cache
online-serving
scheduling
area/ai-infra
2026年8月20日
NSA-ACL25
sparse-attention
long-context
attention-kernel
llm-training
llm-inference
area/ai-infra
2026年8月20日
OD-MoE-arXiv25
llm-inference
moe
edge-inference
expert-loading
distributed-inference
quantization
area/ai-infra
2026年8月20日
PASTA-ICLR24
attention-steering
llm-inference
prompting
model-profiling
inference-time-control
area/ai-infra
2026年8月20日
PhoenixOS-SOSP25
gpu
checkpoint-restore
migration
serverless
fault-tolerance
area/ai-infra
2026年8月20日
PithTrain-arXiv26
moe-training
agent-native
coding-agent
distributed-training
ml-systems
benchmark
area/ai-infra
2026年8月20日
ProfInfer-MLSys26
profiling
ebpf
llm-inference
edge
llama-cpp
observability
area/ai-infra
2026年8月20日
RLBoost-NSDI26
llm-training
reinforcement-learning
spot-instances
rollout
elasticity
area/ai-infra
2026年8月20日
Relax-ASPLOS25
ml-compiler
dynamic-shapes
symbolic-shapes
tvm
deployment
area/ai-infra
2026年8月20日
SAVE-ATC25
fault-tolerance
gpu
inference
edge-ai
bit-flip
area/ai-infra
2026年8月20日
SDCHunter-OSDI26
gpu-reliability
silent-data-corruption
llm-training
deterministic-replay
fault-diagnosis
area/ai-infra
2026年8月20日
SOL-ExecBench-arXiv26
gpu-kernels
benchmark
hardware-roofline
coding-agent
reward-hacking
area/ai-infra
domain/auto-research
2026年8月20日
Sereno-OSDI26
mobile-systems
llm-inference
memory-bandwidth
qos
speculative-decoding
area/ai-infra
2026年8月20日
Sirius-ATC25
gpu-sharing
ml-inference
ml-training
colocation
memory-management
kv-cache
area/ai-infra
2026年8月20日
SkVM-SOSP26
agent-skills
llm-agent
compiler
runtime
jit
area/ai-infra
area/agent-systems
2026年8月20日
SkyServe-EuroSys25
ml-serving
spot-instances
multi-cloud
geo-distributed
availability
area/ai-infra
2026年8月20日
SkyWalker-EuroSys26
llm-inference
multi-region
load-balancing
prefix-caching
cloud-cost
area/ai-infra
2026年8月20日
SolidAttention-FAST26
llm-inference
kv-cache
ssd-offload
attention-sparsity
aipc
area/ai-infra
2026年8月20日
SuperServe-NSDI25
ml-serving
supernet
slo
reactive-scheduling
bursty-workloads
area/ai-infra
2026年8月20日
TapML-ISSTA25
ml-deployment
testing
debugging
webgpu
model-porting
area/ai-infra
2026年8月20日
Tilus-ASPLOS26
gpu-dsl
low-precision
quantization
tensor-layout
llm-serving
area/ai-infra
2026年8月20日
VibeTensor-arXiv26
agent-generated-code
deep-learning-runtime
cuda
autograd
software-engineering
area/ai-infra
2026年8月20日
XGrammar-MLSys25
structured-generation
constrained-decoding
context-free-grammar
llm-serving
agents
area/ai-infra
2026年8月20日
XGrammar2-CAIS26
structured-generation
constrained-decoding
tool-calling
llm-serving
agents
area/ai-infra
area/agent-systems
2026年8月20日
XSched-OSDI25
gpu-scheduling
preemption
xpu
npu
accelerator
area/ai-infra