Remote GPU Server Workflow: Running AI Agents with SSH and Port Forwarding
If you're developing AI models or running inference workloads, you've probably hit the limitations of local hardware. MacBooks can't run CUDA workloads, consumer GPUs lack VRAM for large models, and upgrading to a workstation-class machine is expensive. Remote GPU servers offer a practical alternative: rent powerful hardware by the hour, work from your familiar local environment, and scale up or down as needed.
This guide walks through setting up a productive remote GPU development workflow, from SSH configuration to running AI coding agents on remote hardware.
Why Work on a Remote GPU Server
The case for remote development is straightforward:
Hardware limitations: Apple Silicon Macs don't support CUDA, which most ML frameworks rely on. Even high-end Intel laptops rarely have GPUs suitable for training or running large models.
Cost efficiency: A Lambda Cloud A100 instance costs around $1.10/hour. If you're training models for 20 hours a month, that's $22 versus thousands for purchasing equivalent hardware. For burst workloads—fine-tuning a model over a weekend, running inference experiments, or testing training pipelines—renting is far more economical.
Flexibility: Need 8xA100s for a few hours? Spin up a multi-GPU instance, run your job, and shut it down. No capital expense, no depreciation, no cooling costs.
Environment consistency: Keep your local machine clean and fast. Heavy dependencies, CUDA toolkits, and massive datasets live on the remote server, not cluttering your laptop.
Choosing a GPU Provider
Several providers cater to ML workloads. Here's a factual comparison:
Lambda Cloud offers instances with NVIDIA GPUs (A100, H100, RTX 6000 Ada) at transparent hourly rates. The platform is straightforward—no complex pricing tiers, and instances come with PyTorch and TensorFlow pre-installed. Good documentation, stable network, and predictable availability, though capacity can be tight for popular GPU types.
Vast.ai operates as a marketplace where individuals and datacenters list GPU capacity. You'll find lower prices, especially for older hardware (RTX 3090, RTX 4090), but reliability varies by host. Some machines have residential internet connections, SSH performance can be inconsistent, and you're responsible for OS setup. Best for cost-sensitive workloads where occasional interruptions are acceptable.
RunPod balances price and reliability. They offer both on-demand and spot instances, support persistent storage, and have templates for common ML frameworks. Their network is generally stable, and they support both GPU and CPU-only instances. The interface is more feature-rich than Lambda's, which can be helpful or overwhelming depending on your needs.
Bare metal options: If you're at a company or university, you might have access to an internal GPU cluster. These typically offer better network performance (no cloud overhead) and no metering, but require more setup and maintenance.
For occasional use, Lambda Cloud or RunPod on-demand instances are reliable starting points. For cost-sensitive workloads, Vast.ai spot instances or RunPod spot pricing can cut costs significantly.
Initial SSH Setup
Getting SSH configured properly saves hours of frustration. Start by generating an SSH key if you don't have one:
ssh-keygen -t ed25519 -C "your-email@example.com" -f ~/.ssh/lambda_gpu
Add the public key (~/.ssh/lambda_gpu.pub) to your provider's dashboard. Lambda Cloud and RunPod have SSH key sections in account settings; Vast.ai lets you paste keys when launching instances.
Once your instance is running, note its IP address. Add an entry to ~/.ssh/config:
Host lambda-gpu
HostName 150.136.42.18
User ubuntu
IdentityFile ~/.ssh/lambda_gpu
ServerAliveInterval 60
ServerAliveCountMax 3
Compression yes
The ServerAliveInterval and ServerAliveCountMax options keep the connection alive through network hiccups. Compression yes can speed up transfers on slower connections.
Now you can connect with:
ssh lambda-gpu
ProxyJump for Bastion Hosts
Some providers use bastion hosts for security. If your provider gives you instructions like "SSH to gateway.provider.com, then SSH to your instance," use ProxyJump:
Host gpu-bastion
HostName gateway.provider.com
User your-username
IdentityFile ~/.ssh/provider_key
Host remote-gpu
HostName 10.0.1.50
User ubuntu
IdentityFile ~/.ssh/provider_key
ProxyJump gpu-bastion
This lets you run ssh remote-gpu and SSH handles the two-hop connection automatically.
Essential Port Forwarding Setup
You'll want to access services running on the remote machine from your local browser. Port forwarding maps remote ports to localhost.
Common Ports
- 8888: Jupyter notebooks
- 6006: TensorBoard
- 8000: FastAPI dev servers
- 5000: Flask dev servers
- 7860: Gradio demos
- 8501: Streamlit apps
You can forward ports with the -L flag:
ssh -L 8888:localhost:8888 -L 6006:localhost:6006 lambda-gpu
This gets tedious. Instead, add forwarding rules to your SSH config:
Host lambda-gpu
HostName 150.136.42.18
User ubuntu
IdentityFile ~/.ssh/lambda_gpu
ServerAliveInterval 60
ServerAliveCountMax 3
Compression yes
LocalForward 8888 localhost:8888
LocalForward 6006 localhost:6006
LocalForward 8000 localhost:8000
LocalForward 5000 localhost:5000
Now ssh lambda-gpu automatically forwards these ports. Start Jupyter on the remote machine:
jupyter lab --no-browser --port=8888
Open http://localhost:8888 on your local machine, and you're connected to the remote Jupyter instance.
Dynamic Port Forwarding
If you need flexibility, use dynamic port forwarding with the -D flag:
ssh -D 8080 lambda-gpu
Configure your browser to use localhost:8080 as a SOCKS5 proxy, and all traffic routes through the remote machine. Useful for accessing multiple services without pre-defining ports.
File Synchronization
You need to get code onto the remote machine and retrieve results. Three approaches:
rsync for One-Way Syncs
rsync efficiently copies files, sending only changes. To sync a project directory to the remote server:
rsync -avz --exclude 'venv' --exclude '__pycache__' --exclude '.git' --exclude '*.pyc' --exclude 'data/' ~/Projects/my-model/ lambda-gpu:~/my-model/
Flags:
-a: Archive mode (preserves permissions, timestamps)-v: Verbose output-z: Compress during transfer--exclude: Skip unnecessary files
To sync results back:
rsync -avz lambda-gpu:~/my-model/outputs/ ~/Projects/my-model/outputs/
Watching for Changes
For active development, manually running rsync after every file change is tedious. Use fswatch (macOS) or inotifywait (Linux) to trigger syncs automatically:
fswatch -o ~/Projects/my-model | while read; do
rsync -avz --exclude 'venv' --exclude '__pycache__' --exclude '.git' --exclude '*.pyc' ~/Projects/my-model/ lambda-gpu:~/my-model/
done
This watches for file changes and syncs immediately. For a more robust solution, use watchman or tools like mutagen that handle bidirectional sync and conflict resolution.
When to Use Git Instead
If your project is version-controlled and you're working solo or with a disciplined workflow, git push/pull is simpler:
Local machine:
git add .
git commit -m "Update model architecture"
git push
Remote machine:
git pull
This approach works well for code but poorly for large datasets or model checkpoints (unless you're using Git LFS). For datasets, use rsync or cloud storage (S3, GCS) mounted via tools like rclone or s3fs.
Running AI Coding Agents on the Remote Machine
AI coding agents like Claude Code, Aider, or Codex CLI can write and modify code directly. Running them on the remote machine means they have access to the GPU environment, can test CUDA code immediately, and work with the actual deployment configuration.
Installing Claude Code on the Server
SSH to your remote instance:
ssh lambda-gpu
Install Claude Code (or your preferred agent):
# Using the official installer
curl -fsSL https://download.claudecode.com/install.sh | bash
# Or via npm
npm install -g @anthropic-ai/claude-code
Set your API key:
export ANTHROPIC_API_KEY='your-key-here'
Add this to ~/.bashrc or ~/.zshrc so it persists.
Persistent Sessions with zellij or tmux
SSH connections drop. If you're running a long coding session with an AI agent, you don't want to lose progress when your connection hiccups. Use a terminal multiplexer:
zellij (modern, user-friendly):
# Install
curl -fsSL https://zellij.dev/install.sh | bash
# Start a session
zellij attach -c ai-dev
tmux (traditional, widely available):
# Start a named session
tmux new -s ai-dev
# Detach: Ctrl-b, then d
# Reattach later: tmux attach -t ai-dev
Inside the persistent session, start your AI agent:
claude-code
Now you can interact with the agent, have it write code, run tests, and modify files. If your SSH connection drops, reconnect and reattach:
ssh lambda-gpu
zellij attach ai-dev
Everything is still running.
Using the Agent to Work on GPU Code
A realistic workflow:
# Inside zellij/tmux on remote machine
cd ~/my-model
claude-code
# Agent prompt:
# "Write a PyTorch training script for fine-tuning BERT on a custom dataset.
# Use the Hugging Face transformers library, support multi-GPU training with DDP,
# and log metrics to TensorBoard."
The agent generates the script. You can ask it to refine error handling, adjust hyperparameters, or add checkpointing. Since it's running on the remote machine, it can verify imports, check CUDA availability, and even run test commands:
# Agent prompt:
# "Run a quick test to verify the training script works with a tiny batch."
The agent executes python train.py --batch-size 2 --max-steps 5 and reports success or debugging information.
The Complete Workflow: A Realistic Session
Here's what a typical remote GPU session looks like:
1. SSH In and Attach to Persistent Session
ssh lambda-gpu
zellij attach ai-dev
If this is your first time, create the session:
zellij -c ai-dev
2. Run AI Agent to Write Training Code
cd ~/my-model
claude-code
Prompt the agent:
Write a training script for fine-tuning Llama 3.1 8B on a custom instruction dataset.
Use QLoRA for efficient fine-tuning, log to TensorBoard, save checkpoints every 500 steps.
The agent writes train.py, config.yaml, and a requirements file. Review the code, ask for adjustments if needed.
3. Start Training in Another Pane
Create a new zellij pane (Ctrl-p, then n) or tmux pane (Ctrl-b, then "):
python train.py --config config.yaml
Training starts. GPU utilization hits 95%. You can detach and leave it running.
4. Monitor with TensorBoard
In another pane, start TensorBoard:
tensorboard --logdir ./runs --port 6006
On your local machine, open http://localhost:6006 (port forwarding is already configured in ~/.ssh/config). You see live training metrics: loss curves, learning rate schedules, gradient norms.
5. Sync Results Back
Training finishes. Sync checkpoints and logs to your local machine:
# From local machine
rsync -avz lambda-gpu:~/my-model/checkpoints/ ~/Projects/my-model/checkpoints/
rsync -avz lambda-gpu:~/my-model/runs/ ~/Projects/my-model/runs/
Now you have local copies for further analysis or deployment.
6. Detach and Leave
Detach from zellij (Ctrl-o, then d) or tmux (Ctrl-b, then d). Exit SSH. The session keeps running. If you need to check on it later, SSH back in and reattach.
Cost Optimization Tips
GPU time is metered. A few practices to minimize costs:
Use spot instances: RunPod and Vast.ai offer spot pricing at 50-80% discounts. Your instance can be preempted, but for interruptible workloads (training with checkpointing), the savings are significant.
Auto-shutdown scripts: Create a systemd timer or cron job that shuts down the instance after a period of inactivity:
#!/bin/bash
# /usr/local/bin/auto-shutdown.sh
IDLE_MINUTES=30
UPTIME=$(awk '{print int($1)}' /proc/uptime)
LAST_ACTIVITY=$(last -n 1 -F | awk 'NR==1 {print $NF}')
if [ $((UPTIME - LAST_ACTIVITY)) -gt $((IDLE_MINUTES * 60)) ]; then
sudo shutdown -h now
fi
Run this every 10 minutes via cron:
*/10 * * * * /usr/local/bin/auto-shutdown.sh
Adjust based on your provider's shutdown API.
Monitor GPU usage: Use nvidia-smi or nvitop to check GPU utilization. If you're not using the GPU, you're wasting money. Shut down instances when not actively training or running inference.
Right-size your instance: Don't rent an 8xA100 instance for a workload that fits on a single RTX 4090. Lambda Cloud and RunPod show estimated costs per hour—review them before launching.
Persistent storage: Some providers charge separately for storage. If you're storing large datasets, ensure you're not paying for idle storage when instances are off. Use object storage (S3, GCS) for long-term dataset storage and mount it on-demand.
Integrating Tools for Smoother Workflows
This guide covered the manual setup: SSH config, rsync scripts, port forwarding, and terminal multiplexers. It works, but it's a lot of moving parts.
Tools like Agents UI streamline this workflow by integrating SSH, file transfer, and port forwarding into a native interface. You can manage multiple remote sessions, drag-and-drop files for rsync transfers, and forward ports with UI controls instead of editing config files. For developers who work with remote GPU servers regularly, reducing the friction of context-switching between local editing and remote execution can significantly improve productivity.
Whether you automate with scripts or use an integrated tool, the principles remain the same: reliable SSH configuration, efficient file synchronization, persistent sessions, and cost-conscious resource management.
Conclusion
Remote GPU servers unlock powerful hardware without the capital expense of owning it. With proper SSH setup, port forwarding, file synchronization, and persistent sessions, you can develop and train AI models as if the GPU were local. Running AI coding agents directly on the remote machine further reduces context-switching, letting you iterate faster.
The workflow described here—SSH config with port forwarding, rsync for file sync, zellij or tmux for persistence, and AI agents for code generation—provides a solid foundation for productive remote development. Adjust it to your specific needs, choose providers that match your budget and reliability requirements, and optimize for cost by shutting down idle instances.
Whether you're fine-tuning language models, training computer vision systems, or experimenting with new architectures, remote GPU servers offer a practical path to scalable, cost-effective AI development.