18 KiB
Technical Writeup: OVH Gridscale, Caddy & Prometheus Infrastructure
Overview
This document provides a comprehensive technical writeup of the infrastructure deployment and configuration management workflow. The project provisions cloud resources on OVH / Gridscale using Terraform and automates host configuration using Ansible.
1. Initial Infrastructure Setup & Technical Caveats
1.1 Terraform Provider & Resource Foundation
The initial configuration (main.tf, variables.tf, resources.tf) established authentication with the official Gridscale provider (gridscale/gridscale). Infrastructure resources were defined using the Debian 13 image template and SSH key authentication.
1.2 Key Failures & Solutions
During initial testing and deployment, several technical hurdles were encountered and resolved:
-
SSH Connection Timeouts (Public Network Requirement):
- Issue: Initial attempts to connect via SSH to the server's public IPv4 address timed out (
ssh: connect to host ... port 22: Operation timed out). - Root Cause & Fix: Gridscale instances require explicit attachment to the public network. Because no dynamic network UUID data lookup was provided in the task specification, the public network UUID (
5557a73b-31ee-4b1f-aa15-7789ad6ae04c) was hardcoded directly into the server's network block inresources.tf.
- Issue: Initial attempts to connect via SSH to the server's public IPv4 address timed out (
-
Unpowered Server Instances:
- Issue: Newly created instances remained offline and unreachable after
terraform apply. - Root Cause & Fix: Gridscale creates servers in an unpowered state by default. Setting
power = trueexplicitly ingridscale_serverresources ensured servers automatically boot upon creation.
- Issue: Newly created instances remained offline and unreachable after
-
Ansible SSH Host Key Verification Failures:
- Issue: Ansible fact-gathering failed with
Host key verification failed. - Root Cause & Fix: Cloud IPs change dynamically on fresh deployments. Added
ansible_ssh_extra_args='-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null'to the[all:vars]block ininventory.tfto bypass interactive host key verification prompts seamlessly.
- Issue: Ansible fact-gathering failed with
2. Network Topology & Firewall Security
2.1 Multi-Server Topology
The architecture consists of two dedicated nodes:
caddy01: Web server hosting Caddy.monitoring01: Telemetry node hosting Prometheus.
2.2 Dual-Interface Network Isolation & Inbound Firewall Rules
To secure management endpoints and telemetry data, instances were configured with dual network interfaces:
-
Public Network:
- Attached to Gridscale public network UUID
5557a73b-31ee-4b1f-aa15-7789ad6ae04c. - Firewall Rules (
rules_v4_in,rules_v6_in):- Allow TCP port
22(SSH). - Allow TCP port
80(HTTP). - Block all other inbound public traffic.
- Allow TCP port
- Attached to Gridscale public network UUID
-
Private Network (
gridscale_network.network_internal):- Internal DHCP subnet:
192.168.121.0/27(Gateway:192.168.121.1, DNS:192.168.121.2, Reserved:192.168.121.0/31). - Firewall Rules:
- On
caddy01, allows inbound TCP port2019(Caddy Admin API / Metrics) only from the internal CIDR192.168.121.0/27. Public access to port 2019 is strictly denied at the firewall layer.
- On
- Internal DHCP subnet:
3. Dynamic Ansible Inventory Generation
3.1 Automated Inventory Build (inventory.tf)
To eliminate manual IP tracking, Terraform's local_file resource generates inventory.ini dynamically upon terraform apply.
3.2 Dynamic IP Interpolation
- Extracts auto-assigned private IP addresses from resource state:
${gridscale_server.server_caddy.network[1].auto_assigned_ip}. - Automatically exports host groups (
[caddy],[monitoring]) and internal IP variables (privat_ip_caddy,privat_ip_monitoring) for Ansible playbooks. - Exposes outputs
caddy_public_ip,monitoring_public_ip,caddy_private_ip, andmonitoring_private_ipinoutputs.tf.
4. Ansible Playbook Orchestration (playbook.yml)
4.1 Base System Maintenance & Idempotent Reboot (all)
- Package Updates: Executes
apt updateandupgrade: dist. - Idempotent Reboot Check: Uses
ansible.builtin.statto check for/var/run/reboot-required. The reboot task triggers only when required (when: reboot_required_file.stat.exists == true), avoiding unnecessary reboot delays during repeat runs. - Prerequisites: Installs
debian-keyring,debian-archive-keyring,apt-transport-https, andcurl.
4.2 Caddy Web Server Configuration (caddy)
- Official Deb822 Repository: Per Caddy official documentation, configured Caddy's Debian repository using
ansible.builtin.deb822_repository(https://dl.cloudsmith.io/public/caddy/stable/deb/debianwith official GPG key verification). - Metrics & Admin API Security:
- Configured
Caddyfilewith/metricsand bound the Admin API to:2019. - Set
origins http://{{ privat_ip_caddy }}:2019so Caddy validates origin headers and only processes authorized internal requests.
- Configured
- Service Handlers: Uses Ansible handlers to restart the
caddyservice automatically uponCaddyfileupdates.
4.3 Prometheus Monitoring Setup (monitoring)
- Installation: Installs the
prometheuspackage onmonitoring01. - Private Scrape Target: Configured
prometheus.ymlto scrape Caddy metrics target{{ privat_ip_caddy }}:2019strictly over the internal private network (192.168.121.0/27). - Service Handlers: Triggers
Restart Prometheus serviceautomatically whenprometheus.ymlchanges.
5. Execution Guide
5.1 Provision Infrastructure (Terraform)
# Initialize Terraform providers
terraform init
# Apply infrastructure changes
terraform apply -auto-approve
5.2 Run Ansible Provisioning
# Provision Caddy Web Server
uv run ansible-playbook -i inventory.ini -l caddy playbook.yml
# Provision Monitoring Server
uv run ansible-playbook -i inventory.ini -l monitoring playbook.yml
# Provision All Servers
uv run ansible-playbook -i inventory.ini playbook.yml
6. Modular Jinja2 Templating & Prometheus Alerting Rules
6.1 Refactoring Inline Configurations to Jinja2 Templates
To improve maintainability and separate configuration data from playbook logic, inline file content blocks in playbook.yml were refactored into dedicated Jinja2 template files located in the templates/ directory:
- templates/Caddyfile.j2: Contains Caddy server block definitions and metrics setup.
- templates/prometheus.yml.j2: Contains global scrape configuration and imports rule files via
rule_files: [/etc/prometheus/alert_rules.yml]. - templates/alert_rules.yml.j2: Defines Prometheus alert rules for monitoring Caddy service availability and error rates.
The playbook tasks were updated to use ansible.builtin.template instead of ansible.builtin.copy.
6.2 Variable Naming Refactoring
Standardized variable naming from privat_ip_* to private_ip_* across:
inventory.tf(Terraform dynamic inventory builder)inventory.ini(Generated inventory)playbook.yml& Jinja2 templates ({{ private_ip_caddy }})
6.3 Prometheus Alert Rules & Raw Jinja Escaping
Configured alerting rules in templates/alert_rules.yml.j2 using Jinja {% raw %} ... {% endraw %} blocks to prevent Ansible from misinterpreting Prometheus template variables (such as {{ $labels.instance }} and {{ $value }}):
Caddy Down: Triggers a critical alert ifup{job="caddy"} == 0for 1 minute.CaddyHighHTTP4xxErrorRateService: Triggers a critical alert if the 4xx HTTP error rate exceeds 5% over a 3-minute window (with total request count > 0).- Source: Rule specification adapted from Awesome Prometheus Alerts (Caddy Rules).
6.4 Verification
Verified playbook execution and rule rendering:
uv run ansible-playbook -i inventory.ini -l monitoring playbook.yml
6.5 Automated Playbook Execution via Terraform (null_resource)
To fully automate the end-to-end deployment, a null_resource provisioner was integrated into inventory.tf. It executes automatically as the final step of terraform apply after server creation and inventory file generation:
resource "null_resource" "ansible_provisioner" {
depends_on = [
local_file.ansible_inventory,
gridscale_server.server_caddy,
gridscale_server.server_monitoring
]
provisioner "local-exec" {
command = "sleep 60 &&uv run ansible-playbook -i ${local_file.ansible_inventory.filename} playbook.yml"
}
}
7. End-to-End Environment Setup & Deployment Workflow
7.1 Environment Prerequisites & Credentials (.env)
Copy the .env.example template and export your Gridscale API credentials into your shell session:
# Copy template environment file
cp .env.example .env
# Edit .env with your credentials, then export variables
source .env
7.2 Python Environment Initialization & Dependency Setup (uv)
Initialize Python environment and install Ansible along with passlib and bcrypt (required for Jinja password_hash filter) using uv:
# Initialize uv virtual environment and lockfile
uv init --bare
# Add Ansible, passlib, and bcrypt dependencies (bcrypt<4.0.0 required for passlib compatibility)
uv add ansible passlib "bcrypt<4.0.0"
7.3 Infrastructure Provisioning & Automated Execution
Initialize Terraform and apply the configuration. Terraform will provision the servers, dynamically construct inventory.ini, and automatically trigger the Ansible playbook via the null_resource:
# Initialize Terraform providers
terraform init
# Apply infrastructure changes (automatically provisions servers & runs Ansible)
terraform apply -auto-approve
(Optional) Execute Ansible playbooks manually via uv:
# Run Ansible playbook for all hosts
uv run ansible-playbook -i inventory.ini playbook.yml
# Target specific host groups
uv run ansible-playbook -i inventory.ini -l caddy playbook.yml
uv run ansible-playbook -i inventory.ini -l monitoring playbook.yml
8. Additional Task: Nginx Integration and Custom Improvements
8.1 Nginx Infrastructure & Ansible Deployment
Expanded the infrastructure to provision a dedicated Nginx web server node (nginx01) alongside Caddy and Monitoring nodes:
- Terraform Infrastructure: Defined
gridscale_storage.storage_nginxandgridscale_server.server_nginxattached to the internal private network (gridscale_network.network_internal) with port 80 allowed internally. - Ansible Automation: Added an
nginxplay inplaybook.ymlinstalling Nginx and deploying a customindex.html.j2page displaying host metadata.
8.2 HTTPS Reverse Proxying via sslip.io Dynamic Subdomains
Configured Caddy as the central ingress gateway utilizing sslip.io dynamic DNS wildcard resolution to route traffic over HTTPS automatically:
caddy.<ip>.sslip.io\rightarrowCaddy default static pagenginx.<ip>.sslip.io\rightarrowReverse proxies to private Nginx IP ({{ private_ip_nginx }}:80)monitoring.<ip>.sslip.io\rightarrowReverse proxies to internal Prometheus IP ({{ private_ip_monitoring }}:9090)
8.3 Prometheus Subdomain Security via Basic Authentication & Health Checks
Secured access to the Prometheus monitoring interface by embedding Caddy's basic_auth directive and active health checking directly into the monitoring subdomain block in templates/Caddyfile.j2:
monitoring.{{ ansible_default_ipv4.address | replace(".", "-") }}.sslip.io {
reverse_proxy {{ private_ip_monitoring }}:9090 {
health_uri /-/healthy
health_interval 10s
health_timeout 5s
}
basic_auth {
{{ demo_user }} {{ demo_pass | password_hash('bcrypt') }}
}
}
8.4 Python Dependencies & bcrypt Version Pinning
Encountered and resolved a critical dependency issue during Ansible template rendering for password_hash('bcrypt'):
- Issue:
passlib(1.7.4) failed withAttributeError: module 'bcrypt' has no attribute '__about__'when paired withbcrypt >= 4.0.0, resulting in apassword cannot be longer than 72 bytesexception. - Fix: Pinned
bcryptto<4.0.0(bcrypt==3.2.2) inpyproject.tomland installed viauv:uv add ansible passlib "bcrypt<4.0.0"
8.5 Direct Execution Output via Terraform nonsensitive()
To display all endpoint URLs and auto-generated credentials immediately upon terraform apply completion without manual CLI commands:
- Wrapped
random_password.demo_pass.resultusing Terraform'snonsensitive()function inoutputs.tf:
output "website_urls" {
description = "Website URLs"
value = [
"Caddy: https://caddy.${replace(gridscale_ipv4.public_ipv4_caddy.ip, ".", "-")}.sslip.io",
"Nginx behind Caddy: https://nginx.${replace(gridscale_ipv4.public_ipv4_caddy.ip, ".", "-")}.sslip.io",
"Prometheus behind Caddy: https://monitoring.${replace(gridscale_ipv4.public_ipv4_caddy.ip, ".", "-")}.sslip.io",
"Login for Monitoring: demo ${nonsensitive(random_password.demo_pass.result)}"
]
}
8.6 Upstream Health Checks & Fault Tolerance in Caddy
Configured active health checking inside Caddy's reverse_proxy directives in templates/Caddyfile.j2 to ensure fault tolerance for upstream services:
- Nginx Health Probes: Periodic
health_uri /checks every 10 seconds (5s timeout). - Prometheus Health Probes: Periodic
health_uri /-/healthychecks every 10 seconds (5s timeout). - If an upstream service or node becomes unresponsive, Caddy automatically marks the backend as unhealthy and prevents routing traffic to dead backends.
9. Summary: Service Execution & Access Guide
9.1 Quickstart Execution Steps
-
Configure Environment Credentials: Copy
.env.exampleto.env, set your Gridscale API credentials, and export them:cp .env.example .env # Edit .env with your credentials source .env -
Initialize Python Environment & Dependencies: Initialize
uvenvironment and install Ansible,passlib, andbcrypt:uv init --bare uv add ansible passlib "bcrypt<4.0.0" -
Provision Infrastructure & Execute Playbooks: Apply Terraform to provision all cloud resources, generate
inventory.ini, and automatically trigger Ansible configuration management vianull_resource:terraform init terraform apply -auto-approve
9.2 Service Access & Endpoints Summary
Upon successful completion, Terraform outputs all HTTPS endpoints and Basic Authentication credentials:
| Service | Host / Subdomain Format | Target / Upstream | Access Details |
|---|---|---|---|
| Caddy Ingress | https://caddy.<caddy-ip-with-dashes>.sslip.io |
Local Caddy static server | Public landing page |
| Nginx Web Server | https://nginx.<caddy-ip-with-dashes>.sslip.io |
http://<nginx-ip-private>:80 |
Reverse proxied Hello World page |
| Prometheus Monitoring | https://monitoring.<caddy-ip-with-dashes>.sslip.io |
http://<monitoring-ip-private>:9090 with caddy basic_auth |
Username: demoPassword: Auto-generated password from Terraform output |
(Optional) Display all endpoint URLs and active credentials at any time:
terraform output website_urls
10. Exercise Requirement Traceability & Execution Summary
The following matrix maps every requirement, deliverable, and bonus task from the Gridscale hiring exercise prompt to its exact technical implementation, file location, and execution method in this repository:
| Exercise Requirement / Task | Status | Implementation Details & File Location | Execution & Automation Method |
|---|---|---|---|
| Install & Configure Caddy Server | Completed | Provisioned caddy01 VM on Debian 13 in resources.tf. Installed Caddy via Deb822 repo in playbook.yml. |
Automated via Ansible ansible.builtin.apt & ansible.builtin.template |
| Serve Sample Web Application | Completed | Configured static file server (root * /usr/share/caddy, file_server) in templates/Caddyfile.j2. |
Rendered via Ansible template Caddyfile.j2 |
| Prometheus Monitoring Setup | Completed | Provisioned monitoring01 VM in resources.tf. Installed Prometheus in playbook.yml. |
Configured scrape job caddy targeting {{ private_ip_caddy }}:2019 |
| Metrics Collection (Codes, Latency, Uptime) | Completed | Enabled Caddy metrics directive metrics /metrics with per_host and observe_catchall_hosts on Admin API :2019 in templates/Caddyfile.j2. |
Scraped every 15s by Prometheus over internal network (192.168.121.0/27) |
| Prometheus Alerting Pipeline & Rules | Completed | Defined Caddy Down (up{job="caddy"} == 0) and CaddyHighHTTP4xxErrorRateService (> 5% 4xx rate) rules in templates/alert_rules.yml.j2. Imported in templates/prometheus.yml.j2. |
Deployed to /etc/prometheus/alert_rules.yml via Ansible |
| IaC Automation (Terraform + Ansible) | Completed | 100% automated with Terraform (Gridscale provider) & Ansible. Built null_resource.ansible_provisioner in inventory.tf to execute Ansible automatically. |
Runs end-to-end via terraform apply -auto-approve |
| Additional Task: Nginx "Hello World" VM | Completed | Provisioned nginx01 VM in resources.tf. Installed Nginx and deployed index.html.j2 in playbook.yml. |
Bound strictly to internal private IP (192.168.121.5:80) |
| Additional Task: Caddy Reverse Proxy & Routing | Completed | Configured subdomain reverse proxy routing in templates/Caddyfile.j2: nginx.<ip>.sslip.io \rightarrow Nginx, monitoring.<ip>.sslip.io \rightarrow Prometheus. |
Managed via Caddy ingress rules |
| Bonus: Upstream Health Checks & Fault Tolerance | Completed | Added health_uri / (Nginx) and health_uri /-/healthy (Prometheus) with 10s intervals in templates/Caddyfile.j2. |
Probed actively by Caddy to bypass dead nodes |
| Bonus: SSL Termination & Client Encryption | Completed | Used sslip.io dynamic wildcard subdomains on port 443 in templates/Caddyfile.j2. |
Caddy Auto-HTTPS manages Let's Encrypt TLS termination |
| Deliverables: Endpoints & Output Credentials | Completed | Configured unmasked output using nonsensitive() in outputs.tf. |
Printed directly to terminal on terraform apply |