> ## Documentation Index
> Fetch the complete documentation index at: https://docs.arc.computer/llms.txt
> Use this file to discover all available pages before exploring further.

# Frequently Asked Questions

> Common questions about ATLAS implementation and usage

## General Questions

### What is ATLAS?

ATLAS (Adaptive Teaching and Learning Alignment System) is a framework that pairs your existing agent (the student) with a specialized verifying teacher that provides adaptive guidance. It uses a two-pass protocol: diagnostic assessment followed by targeted teaching.

### Does ATLAS train on my production data?

No. The runtime loop operates through inference-time feedback—no model weights are modified during production execution. Weight updates only happen when you explicitly run offline GRPO training jobs on your own infrastructure.

**Data control:**

* All session traces write exclusively to the Postgres database you provide via `storage.database_url`
* ATLAS never operates its own data store or accesses your database
* Leave `storage: null` in your config to run in ephemeral mode with no persistent data
* Training is opt-in: you choose when to export traces and run offline training

The runtime learns by storing successful guidance patterns in memory (when storage is enabled) and retrieving them for similar future tasks—this is inference-time adaptation, not model training.

### How is ATLAS different from fine-tuning?

Unlike fine-tuning which modifies model weights, ATLAS:

* Preserves the student model's original capabilities
* Works with any model without retraining
* Adapts guidance based on student capability
* Provides immediate enhancement without training time

### What performance improvements can I expect?

Across our evaluation suite we consistently see the closed-loop dual-agent runtime (student + verifying teacher) deliver an average **+15.7% accuracy gain**, **31% task completion lift**, **97% non-degradation**, and **\~50% token savings** versus baseline agents. Offline GRPO training then compounds those gains when you fine-tune custom teacher checkpoints using the traces exported from production.

Actual results vary with task difficulty, data quality, and the strength of the underlying student model, but the closed loop plus GRPO stack gives you levers to reach those numbers. Online continual learning lives in the [`atlas-sdk`](https://github.com/Arc-Computer/atlas-sdk) runtime if you need rapid, task-specific adaptation.

## Hardware & Setup

### What hardware do I need?

**Minimum Requirements:**

* GPU: 16GB VRAM (RTX 4080, A5000)
* RAM: 32GB system memory
* Storage: 100GB for models and data

**Recommended for Training:**

* GPU: 4× A100 40GB or H100 80GB
* RAM: 128GB+ system memory
* Storage: 500GB NVMe SSD

**For Inference Only:**

* Can run on CPU (slower)
* 8GB VRAM with quantization
* Cloud instances work well

### Can I run ATLAS on CPU?

Yes, with API-based models you need no GPU at all. For local model inference:

* CPU inference is 10-50x slower than GPU
* Limited to smaller models (4B-8B)
* Quantization recommended
* Suitable for development/testing

```python theme={null}
# API-based (no GPU needed)
from openai import OpenAI
client = OpenAI()
# Call the verifying teacher and student via API

# For local models on CPU
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
    "Arc-Intelligence/ATLAS-8B-Thinking",
    device_map="cpu",
    torch_dtype=torch.float32
)
```

### Which models are compatible?

**Verifying teacher checkpoints (pre-trained):**

* ATLAS-8B-Thinking (reasoning)
* ATLAS-8B-Instruct (coding)

**Student agents (any LLM):**

* Qwen series (4B-70B)
* Llama series (7B-70B)
* Mistral/Mixtral models
* GPT-3.5/4 (via API)
* Claude (via API)

## Training Questions

### How long does training take?

**Offline RL Training (GRPO):**

* SFT warmup: 4-8 hours
* GRPO training: 24-48 hours
* Hardware: 4-8 H100 GPUs

### What's the difference between online and offline training?

**Offline Training (GRPO):**

* Creates foundational teaching skills
* Requires significant compute
* Produces generalizable models
* One-time investment

**Continual Learning (atlas-sdk):**

* Adapts to specific tasks using the runtime loop
* Runs through the SDK CLI and APIs
* Rapid iteration cycles driven by live traces
* Keeps production agents improving between offline training runs

### Can I train on custom data?

Yes, prepare your data in this format:

```json theme={null}
{
  "prompt": "Your task or question",
  "ground_truth": "Correct answer",
  "metadata": {
    "domain": "your_domain",
    "difficulty": "easy|medium|hard"
  }
}
```

Then train:

```bash theme={null}
scripts/launch.sh 8 src/atlas_core/configs/recipe/teacher_sft.yaml \
  dataset_name=path/to/your/data
```

## Implementation Questions

### How do I integrate ATLAS into my application?

Use the ATLAS teaching protocol with real imports:

```python theme={null}
from openai import OpenAI
from atlas_core.reward.interpretation import RIMReward

client = OpenAI()

# Step 1: Get baseline from your student model
baseline_response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": task}]
).choices[0].message.content

# Step 2: Get teaching from teacher model
teaching = client.chat.completions.create(
    model="gpt-4o",
    messages=[{
        "role": "user",
        "content": f"Provide teaching for: {task}\nStudent said: {baseline_response}"
    }]
).choices[0].message.content

# Step 3: Student applies teaching
enhanced_response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": f"{task}\nTeaching: {teaching}"}]
).choices[0].message.content
```

See [Inference Integration Guide](/sdk/quickstart) for complete examples.

### Can ATLAS work with my existing agent?

Yes. Use the [`atlas-sdk`](https://github.com/Arc-Computer/atlas-sdk) runtime wrappers (HTTP, Python callable, OpenAI Assistants, CLI) to orchestrate your agent, export traces, and hand them to Atlas Core for training. The SDK documentation covers the available adapters and configuration options.

### How do I monitor performance in production?

Use RIM reward scoring to track quality:

```python theme={null}
from atlas_core.reward.interpretation import RIMReward

reward = RIMReward(config_path='reward_system/interpretation.yaml')

# Track each request
baseline_eval = reward.evaluate(prompt=task, response=baseline_response)
enhanced_eval = reward.evaluate(
    prompt=task,
    response=enhanced_response,
    baseline_solutions=baseline_response,
    teacher_traces=teaching
)

# Log metrics
print(f"Baseline: {baseline_eval.score:.3f}")
print(f"Enhanced: {enhanced_eval.score:.3f}")
print(f"Delta: {enhanced_eval.score - baseline_eval.score:+.3f}")
```

RIM scores can be logged to:

* Weights & Biases
* TensorBoard
* Prometheus
* Custom logging systems

## Performance & Optimization

### Why is inference slow?

Common causes and solutions:

1. **Not using Flash Attention**:
   ```python theme={null}
   config.attn_implementation = "flash_attention_2"
   ```

2. **Small batch size**:
   ```python theme={null}
   atlas.batch_size = 8  # Process multiple requests
   ```

3. **No caching**:
   ```python theme={null}
   atlas.enable_cache = True
   ```

4. **CPU inference**: Use GPU or quantization

### How can I reduce memory usage?

Progressive solutions:

1. **Quantization** (75% reduction):
   ```python theme={null}
   config.load_in_4bit = True
   ```

2. **Smaller models**: Use 4B instead of 8B

3. **Offloading**: Move to CPU/disk

4. **Batch size**: Reduce to 1

### What if the teacher makes things worse?

ATLAS has a 97% non-degradation guarantee through:

* Zero reward for performance drops
* Safety validation before deployment
* Fallback to baseline response
* Continuous monitoring

If issues persist:

* Check task-model compatibility
* Verify data quality
* Adjust teaching parameters
* Export fresh traces and schedule a GRPO training run

## Cost Questions

### How much does ATLAS cost to run?

**Training Costs:**

* Offline RL: \$100-500 in compute (depends on GPUs and run length)

**Inference Costs:**

* Self-hosted: Electricity only
* Cloud GPU: \$1-3/hour
* API-based: \$0.001-0.01 per request

### Is there a cloud service?

Currently ATLAS is open-source only. You can:

* Self-host on your infrastructure
* Use cloud GPU providers
* Deploy on Hugging Face Spaces
* Contact team for enterprise support

## Troubleshooting

### Where can I get help?

1. [Troubleshooting Guide](/reference/troubleshooting)
2. [GitHub Issues](https://github.com/Arc-Computer/ATLAS/issues)
3. [Discord Community](https://discord.gg/arc-atlas)
4. Email: [support@arc.computer](mailto:support@arc.computer)

### How do I report a bug?

File an issue with:

* Error message and stack trace
* System configuration
* Minimal reproduction code
* Expected vs actual behavior

### Can I contribute to ATLAS?

Yes! We welcome contributions:

* Code improvements
* Documentation
* Bug fixes
* New features
* Dataset contributions

See [Contributing Guide](https://github.com/Arc-Computer/ATLAS/blob/main/CONTRIBUTING.md).

## Next Steps

<CardGroup cols="2">
  <Card title="Quickstart" icon="rocket" href="/sdk/quickstart">
    Get started with ATLAS
  </Card>

  <Card title="Examples" icon="lightbulb" href="/examples/adaptive-tool-use">
    See ATLAS in action
  </Card>

  <Card title="Troubleshooting" icon="wrench" href="/reference/troubleshooting">
    Solve common issues
  </Card>

  <Card title="Community" icon="discord" href="https://discord.gg/arc-atlas">
    Join the discussion
  </Card>
</CardGroup>
