AI term
Pipeline Parallelism
What is Pipeline Parallelism?
A technique that splits an AI model's layers across multiple processors, allowing different parts of a sequence to be processed simultaneously.
In other languages
- 한국어파이프라인 병렬화
- 딥러닝 모델의 레이어를 여러 개의 GPU에 나누어 할당하고 데이터를 순차적으로 처리하는 병렬 연산 기법
- 日本語パイプライン並列化
- モデルの各層を複数のGPUに分割して割り当て、リレー形式で並列処理を行う手法。
Related Terms
- Expert ParallelismA parallel computing strategy where different components of a neural network are distributed across multiple processors to handle models with parameters too large for a single device.
- Tensor ParallelismA technique for splitting individual model layers across multiple GPU devices to accelerate large-scale model inference and training
- Spatial parallelismParallelization method that divides spatial data regions across multiple processing devices
- Parallel ReasoningAn inference paradigm used in artificial intelligence, especially with LLMs, that involves exploring and processing multiple lines of thought concurrently to solve complex problems more robustly and efficiently than traditional sequential methods.
- Distributed trainingAn approach to machine learning where the computational workload of training a model is split across multiple separate devices or processors.
- Layerwise offloadA technique that moves model layers between GPU and CPU memory as needed to reduce GPU memory usage during inference or training.
- Operator FusionAn optimization technique that combines multiple consecutive neural network operations into a single kernel to reduce memory overhead and improve speed.
- SIMTParallel execution model where many threads run the same instruction on different data
- Object MultiplexingA computational technique where multiple independent data streams or objects are processed simultaneously in a single operation.
- Forward passA computation that carries input data through a neural network to produce an output.
- Neural Processing UnitA specialized hardware accelerator designed specifically to handle the mathematical computations required for machine learning and artificial intelligence tasks with high efficiency.
- Asynchronous TrainingA distributed computing technique where independent processors update a global model at different times without requiring constant, uniform synchronization