Files
ovh_gridscale_test_task/README.md
T

18 KiB

Technical Writeup: OVH Gridscale, Caddy & Prometheus Infrastructure

Overview

This document provides a comprehensive technical writeup of the infrastructure deployment and configuration management workflow. The project provisions cloud resources on OVH / Gridscale using Terraform and automates host configuration using Ansible.


1. Initial Infrastructure Setup & Technical Caveats

1.1 Terraform Provider & Resource Foundation

The initial configuration (main.tf, variables.tf, resources.tf) established authentication with the official Gridscale provider (gridscale/gridscale). Infrastructure resources were defined using the Debian 13 image template and SSH key authentication.

1.2 Key Failures & Solutions

During initial testing and deployment, several technical hurdles were encountered and resolved:

  1. SSH Connection Timeouts (Public Network Requirement):

    • Issue: Initial attempts to connect via SSH to the server's public IPv4 address timed out (ssh: connect to host ... port 22: Operation timed out).
    • Root Cause & Fix: Gridscale instances require explicit attachment to the public network. Because no dynamic network UUID data lookup was provided in the task specification, the public network UUID (5557a73b-31ee-4b1f-aa15-7789ad6ae04c) was hardcoded directly into the server's network block in resources.tf.
  2. Unpowered Server Instances:

    • Issue: Newly created instances remained offline and unreachable after terraform apply.
    • Root Cause & Fix: Gridscale creates servers in an unpowered state by default. Setting power = true explicitly in gridscale_server resources ensured servers automatically boot upon creation.
  3. Ansible SSH Host Key Verification Failures:

    • Issue: Ansible fact-gathering failed with Host key verification failed.
    • Root Cause & Fix: Cloud IPs change dynamically on fresh deployments. Added ansible_ssh_extra_args='-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null' to the [all:vars] block in inventory.tf to bypass interactive host key verification prompts seamlessly.

2. Network Topology & Firewall Security

2.1 Multi-Server Topology

The architecture consists of two dedicated nodes:

  • caddy01: Web server hosting Caddy.
  • monitoring01: Telemetry node hosting Prometheus.

2.2 Dual-Interface Network Isolation & Inbound Firewall Rules

To secure management endpoints and telemetry data, instances were configured with dual network interfaces:

  1. Public Network:

    • Attached to Gridscale public network UUID 5557a73b-31ee-4b1f-aa15-7789ad6ae04c.
    • Firewall Rules (rules_v4_in, rules_v6_in):
      • Allow TCP port 22 (SSH).
      • Allow TCP port 80 (HTTP).
      • Block all other inbound public traffic.
  2. Private Network (gridscale_network.network_internal):

    • Internal DHCP subnet: 192.168.121.0/27 (Gateway: 192.168.121.1, DNS: 192.168.121.2, Reserved: 192.168.121.0/31).
    • Firewall Rules:
      • On caddy01, allows inbound TCP port 2019 (Caddy Admin API / Metrics) only from the internal CIDR 192.168.121.0/27. Public access to port 2019 is strictly denied at the firewall layer.

3. Dynamic Ansible Inventory Generation

3.1 Automated Inventory Build (inventory.tf)

To eliminate manual IP tracking, Terraform's local_file resource generates inventory.ini dynamically upon terraform apply.

3.2 Dynamic IP Interpolation

  • Extracts auto-assigned private IP addresses from resource state: ${gridscale_server.server_caddy.network[1].auto_assigned_ip}.
  • Automatically exports host groups ([caddy], [monitoring]) and internal IP variables (privat_ip_caddy, privat_ip_monitoring) for Ansible playbooks.
  • Exposes outputs caddy_public_ip, monitoring_public_ip, caddy_private_ip, and monitoring_private_ip in outputs.tf.

4. Ansible Playbook Orchestration (playbook.yml)

4.1 Base System Maintenance & Idempotent Reboot (all)

  • Package Updates: Executes apt update and upgrade: dist.
  • Idempotent Reboot Check: Uses ansible.builtin.stat to check for /var/run/reboot-required. The reboot task triggers only when required (when: reboot_required_file.stat.exists == true), avoiding unnecessary reboot delays during repeat runs.
  • Prerequisites: Installs debian-keyring, debian-archive-keyring, apt-transport-https, and curl.

4.2 Caddy Web Server Configuration (caddy)

  • Official Deb822 Repository: Per Caddy official documentation, configured Caddy's Debian repository using ansible.builtin.deb822_repository (https://dl.cloudsmith.io/public/caddy/stable/deb/debian with official GPG key verification).
  • Metrics & Admin API Security:
    • Configured Caddyfile with /metrics and bound the Admin API to :2019.
    • Set origins http://{{ privat_ip_caddy }}:2019 so Caddy validates origin headers and only processes authorized internal requests.
  • Service Handlers: Uses Ansible handlers to restart the caddy service automatically upon Caddyfile updates.

4.3 Prometheus Monitoring Setup (monitoring)

  • Installation: Installs the prometheus package on monitoring01.
  • Private Scrape Target: Configured prometheus.yml to scrape Caddy metrics target {{ privat_ip_caddy }}:2019 strictly over the internal private network (192.168.121.0/27).
  • Service Handlers: Triggers Restart Prometheus service automatically when prometheus.yml changes.

5. Execution Guide

5.1 Provision Infrastructure (Terraform)

# Initialize Terraform providers
terraform init

# Apply infrastructure changes
terraform apply -auto-approve

5.2 Run Ansible Provisioning

# Provision Caddy Web Server
uv run ansible-playbook -i inventory.ini -l caddy playbook.yml

# Provision Monitoring Server
uv run ansible-playbook -i inventory.ini -l monitoring playbook.yml

# Provision All Servers
uv run ansible-playbook -i inventory.ini playbook.yml

6. Modular Jinja2 Templating & Prometheus Alerting Rules

6.1 Refactoring Inline Configurations to Jinja2 Templates

To improve maintainability and separate configuration data from playbook logic, inline file content blocks in playbook.yml were refactored into dedicated Jinja2 template files located in the templates/ directory:

The playbook tasks were updated to use ansible.builtin.template instead of ansible.builtin.copy.

6.2 Variable Naming Refactoring

Standardized variable naming from privat_ip_* to private_ip_* across:

  • inventory.tf (Terraform dynamic inventory builder)
  • inventory.ini (Generated inventory)
  • playbook.yml & Jinja2 templates ({{ private_ip_caddy }})

6.3 Prometheus Alert Rules & Raw Jinja Escaping

Configured alerting rules in templates/alert_rules.yml.j2 using Jinja {% raw %} ... {% endraw %} blocks to prevent Ansible from misinterpreting Prometheus template variables (such as {{ $labels.instance }} and {{ $value }}):

  1. Caddy Down: Triggers a critical alert if up{job="caddy"} == 0 for 1 minute.
  2. CaddyHighHTTP4xxErrorRateService: Triggers a critical alert if the 4xx HTTP error rate exceeds 5% over a 3-minute window (with total request count > 0).

6.4 Verification

Verified playbook execution and rule rendering:

uv run ansible-playbook -i inventory.ini -l monitoring playbook.yml

6.5 Automated Playbook Execution via Terraform (null_resource)

To fully automate the end-to-end deployment, a null_resource provisioner was integrated into inventory.tf. It executes automatically as the final step of terraform apply after server creation and inventory file generation:

resource "null_resource" "ansible_provisioner" {
  depends_on = [
    local_file.ansible_inventory,
    gridscale_server.server_caddy,
    gridscale_server.server_monitoring
  ]

  provisioner "local-exec" {
    command = "sleep 60 &&uv run ansible-playbook -i ${local_file.ansible_inventory.filename} playbook.yml"
  }
}

7. End-to-End Environment Setup & Deployment Workflow

7.1 Environment Prerequisites & Credentials (.env)

Copy the .env.example template and export your Gridscale API credentials into your shell session:

# Copy template environment file
cp .env.example .env

# Edit .env with your credentials, then export variables
source .env

7.2 Python Environment Initialization & Dependency Setup (uv)

Initialize Python environment and install Ansible along with passlib and bcrypt (required for Jinja password_hash filter) using uv:

# Initialize uv virtual environment and lockfile
uv init --bare

# Add Ansible, passlib, and bcrypt dependencies (bcrypt<4.0.0 required for passlib compatibility)
uv add ansible passlib "bcrypt<4.0.0"

7.3 Infrastructure Provisioning & Automated Execution

Initialize Terraform and apply the configuration. Terraform will provision the servers, dynamically construct inventory.ini, and automatically trigger the Ansible playbook via the null_resource:

# Initialize Terraform providers
terraform init

# Apply infrastructure changes (automatically provisions servers & runs Ansible)
terraform apply -auto-approve

(Optional) Execute Ansible playbooks manually via uv:

# Run Ansible playbook for all hosts
uv run ansible-playbook -i inventory.ini playbook.yml

# Target specific host groups
uv run ansible-playbook -i inventory.ini -l caddy playbook.yml
uv run ansible-playbook -i inventory.ini -l monitoring playbook.yml

8. Additional Task: Nginx Integration and Custom Improvements

8.1 Nginx Infrastructure & Ansible Deployment

Expanded the infrastructure to provision a dedicated Nginx web server node (nginx01) alongside Caddy and Monitoring nodes:

  • Terraform Infrastructure: Defined gridscale_storage.storage_nginx and gridscale_server.server_nginx attached to the internal private network (gridscale_network.network_internal) with port 80 allowed internally.
  • Ansible Automation: Added an nginx play in playbook.yml installing Nginx and deploying a custom index.html.j2 page displaying host metadata.

8.2 HTTPS Reverse Proxying via sslip.io Dynamic Subdomains

Configured Caddy as the central ingress gateway utilizing sslip.io dynamic DNS wildcard resolution to route traffic over HTTPS automatically:

  • caddy.<ip>.sslip.io \rightarrow Caddy default static page
  • nginx.<ip>.sslip.io \rightarrow Reverse proxies to private Nginx IP ({{ private_ip_nginx }}:80)
  • monitoring.<ip>.sslip.io \rightarrow Reverse proxies to internal Prometheus IP ({{ private_ip_monitoring }}:9090)

8.3 Prometheus Subdomain Security via Basic Authentication & Health Checks

Secured access to the Prometheus monitoring interface by embedding Caddy's basic_auth directive and active health checking directly into the monitoring subdomain block in templates/Caddyfile.j2:

monitoring.{{ ansible_default_ipv4.address | replace(".", "-") }}.sslip.io {
  reverse_proxy {{ private_ip_monitoring }}:9090 {
    health_uri /-/healthy
    health_interval 10s
    health_timeout 5s
  }
  basic_auth {
    {{ demo_user }} {{ demo_pass | password_hash('bcrypt') }}
  }
}

8.4 Python Dependencies & bcrypt Version Pinning

Encountered and resolved a critical dependency issue during Ansible template rendering for password_hash('bcrypt'):

  • Issue: passlib (1.7.4) failed with AttributeError: module 'bcrypt' has no attribute '__about__' when paired with bcrypt >= 4.0.0, resulting in a password cannot be longer than 72 bytes exception.
  • Fix: Pinned bcrypt to <4.0.0 (bcrypt==3.2.2) in pyproject.toml and installed via uv:
    uv add ansible passlib "bcrypt<4.0.0"
    

8.5 Direct Execution Output via Terraform nonsensitive()

To display all endpoint URLs and auto-generated credentials immediately upon terraform apply completion without manual CLI commands:

  • Wrapped random_password.demo_pass.result using Terraform's nonsensitive() function in outputs.tf:
output "website_urls" {
  description = "Website URLs"
  value = [
    "Caddy: https://caddy.${replace(gridscale_ipv4.public_ipv4_caddy.ip, ".", "-")}.sslip.io",
    "Nginx behind Caddy: https://nginx.${replace(gridscale_ipv4.public_ipv4_caddy.ip, ".", "-")}.sslip.io",
    "Prometheus behind Caddy: https://monitoring.${replace(gridscale_ipv4.public_ipv4_caddy.ip, ".", "-")}.sslip.io",
    "Login for Monitoring: demo ${nonsensitive(random_password.demo_pass.result)}"
  ]
}

8.6 Upstream Health Checks & Fault Tolerance in Caddy

Configured active health checking inside Caddy's reverse_proxy directives in templates/Caddyfile.j2 to ensure fault tolerance for upstream services:

  • Nginx Health Probes: Periodic health_uri / checks every 10 seconds (5s timeout).
  • Prometheus Health Probes: Periodic health_uri /-/healthy checks every 10 seconds (5s timeout).
  • If an upstream service or node becomes unresponsive, Caddy automatically marks the backend as unhealthy and prevents routing traffic to dead backends.

9. Summary: Service Execution & Access Guide

9.1 Quickstart Execution Steps

  1. Configure Environment Credentials: Copy .env.example to .env, set your Gridscale API credentials, and export them:

    cp .env.example .env
    # Edit .env with your credentials
    source .env
    
  2. Initialize Python Environment & Dependencies: Initialize uv environment and install Ansible, passlib, and bcrypt:

    uv init --bare
    uv add ansible passlib "bcrypt<4.0.0"
    
  3. Provision Infrastructure & Execute Playbooks: Apply Terraform to provision all cloud resources, generate inventory.ini, and automatically trigger Ansible configuration management via null_resource:

    terraform init
    terraform apply -auto-approve
    

9.2 Service Access & Endpoints Summary

Upon successful completion, Terraform outputs all HTTPS endpoints and Basic Authentication credentials:

Service Host / Subdomain Format Target / Upstream Access Details
Caddy Ingress https://caddy.<caddy-ip-with-dashes>.sslip.io Local Caddy static server Public landing page
Nginx Web Server https://nginx.<caddy-ip-with-dashes>.sslip.io http://<nginx-ip-private>:80 Reverse proxied Hello World page
Prometheus Monitoring https://monitoring.<caddy-ip-with-dashes>.sslip.io http://<monitoring-ip-private>:9090 with caddy basic_auth Username: demo
Password: Auto-generated password from Terraform output

(Optional) Display all endpoint URLs and active credentials at any time:

terraform output website_urls

10. Exercise Requirement Traceability & Execution Summary

The following matrix maps every requirement, deliverable, and bonus task from the Gridscale hiring exercise prompt to its exact technical implementation, file location, and execution method in this repository:

Exercise Requirement / Task Status Implementation Details & File Location Execution & Automation Method
Install & Configure Caddy Server Completed Provisioned caddy01 VM on Debian 13 in resources.tf. Installed Caddy via Deb822 repo in playbook.yml. Automated via Ansible ansible.builtin.apt & ansible.builtin.template
Serve Sample Web Application Completed Configured static file server (root * /usr/share/caddy, file_server) in templates/Caddyfile.j2. Rendered via Ansible template Caddyfile.j2
Prometheus Monitoring Setup Completed Provisioned monitoring01 VM in resources.tf. Installed Prometheus in playbook.yml. Configured scrape job caddy targeting {{ private_ip_caddy }}:2019
Metrics Collection (Codes, Latency, Uptime) Completed Enabled Caddy metrics directive metrics /metrics with per_host and observe_catchall_hosts on Admin API :2019 in templates/Caddyfile.j2. Scraped every 15s by Prometheus over internal network (192.168.121.0/27)
Prometheus Alerting Pipeline & Rules Completed Defined Caddy Down (up{job="caddy"} == 0) and CaddyHighHTTP4xxErrorRateService (> 5% 4xx rate) rules in templates/alert_rules.yml.j2. Imported in templates/prometheus.yml.j2. Deployed to /etc/prometheus/alert_rules.yml via Ansible
IaC Automation (Terraform + Ansible) Completed 100% automated with Terraform (Gridscale provider) & Ansible. Built null_resource.ansible_provisioner in inventory.tf to execute Ansible automatically. Runs end-to-end via terraform apply -auto-approve
Additional Task: Nginx "Hello World" VM Completed Provisioned nginx01 VM in resources.tf. Installed Nginx and deployed index.html.j2 in playbook.yml. Bound strictly to internal private IP (192.168.121.5:80)
Additional Task: Caddy Reverse Proxy & Routing Completed Configured subdomain reverse proxy routing in templates/Caddyfile.j2: nginx.<ip>.sslip.io \rightarrow Nginx, monitoring.<ip>.sslip.io \rightarrow Prometheus. Managed via Caddy ingress rules
Bonus: Upstream Health Checks & Fault Tolerance Completed Added health_uri / (Nginx) and health_uri /-/healthy (Prometheus) with 10s intervals in templates/Caddyfile.j2. Probed actively by Caddy to bypass dead nodes
Bonus: SSL Termination & Client Encryption Completed Used sslip.io dynamic wildcard subdomains on port 443 in templates/Caddyfile.j2. Caddy Auto-HTTPS manages Let's Encrypt TLS termination
Deliverables: Endpoints & Output Credentials Completed Configured unmasked output using nonsensitive() in outputs.tf. Printed directly to terminal on terraform apply