Repository navigation
database-monitoring: run the deployment updater on Fargate as a second ECS service (module 0.3.0) - #129
Conversation
…d ECS service (module 0.3.0) New inputs updater, updater_image, update_channel and agent_id. With updater = true the module runs elastio-database-monitoring-updater as a second service in the agent's cluster, with its own execution role (reads only the agent's API-key secret) and a task role that may only describe and update the agent's service (to a revision of the agent's family), register that family, describe task definitions, and pass the agent's two roles to ECS. updater requires runtime_updates. Closes #128
| variable "agent_id" { | ||
| description = "The agent's ID in Elastio, given to the updater. Optional: when empty, the updater asks Elastio for it with the agent's API key." | ||
| type = string | ||
| default = "" |
There was a problem hiding this comment.
agent_id must be required when updater = true.
The updater does not ask Elastio for the id. In update-agent.py, an empty ELASTIO_DBMON_AGENT_ID only logs "Updates are off until then" every 5 min.
So with the default "" the updater service runs and costs money, but never updates the agent.
Please add a validation like the runtime_updates one, and fix the agent_id description, the README and the PR body.
There was a problem hiding this comment.
Fixed in 41b2b2b: agent_id is now validated like runtime_updates — updater = true without a UUID agent_id is refused at plan (a non-UUID too; it stays optional with the updater off). Description, README (inputs table regenerated, example and prose) and the PR body are corrected; they now point at where the ID is found: the agent logs registered as <id> on every start in the module's log group. Three new tests (missing, not a UUID, optional when off); 27/27 pass, and neutering the validation fails updater_requires_agent_id.
The updater cannot discover the agent's ID: with an empty one it runs, costs money and never updates the agent (review on #129). Validate it like runtime_updates, point the description and README at where the ID is found (the agent logs "registered as <id>" on every start), and test the refusal, a non-UUID, and that it stays optional with the updater off.
maksvet
left a comment
There was a problem hiding this comment.
LGTM. The agent_id fix is good.
Not blocking: the new .terraform.lock.hcl at the module root can go. In this repo only examples/ keep lock files.
Closes #128.
Fargate deployments have no host to run the deployment updater on, so "Update to latest" never appears for them. With
updater = truethe module now runs the updater as a second ECS service in the agent's cluster.New inputs:
updater(defaultfalse),updater_image(default by channel:public.ecr.aws/elastio/elastio-database-monitoring-updater:0.1.10for production,public.ecr.aws/elastio-development/elastio-database-monitoring-updater:latestfor development),update_channel(production|development),agent_id(default"", required and validated as a UUID whenupdater = true: the updater cannot discover it, and without it the service would run and never update anything; the agent logs it on every start asregistered as <id>). New outputs:updater_service_nameandupdater_task_role_arn. The module version goes to 0.3.0. That release also carriesruntime_updatesfrom #127, which merged after 0.2.1 was tagged.What
updater = truecreates: one service,<name>-updater(desired 1, minimum healthy 0 / maximum 100 so two never run at once), with a 0.25 vCPU / 512 MiB ARM64 task. It runs in the agent's subnets and security groups and needs no inbound rules. It logs to the module's log group underupdater/. Its environment isELASTIO_DBMON_SERVER_URL,_AGENT_ID,_DEPLOYMENT=ecs,_UPDATE_CHANNEL,_ECS_CLUSTER,_ECS_SERVICEand_ECS_CONTAINER, plusELASTIO_DBMON_API_KEYfrom the agent's own secret.The service also gets its own two roles:
AmazonECSTaskExecutionRolePolicy, the same as the agent's (image pull and log write), plussecretsmanager:GetSecretValueon the agent's API-key secret only.ecs:DescribeServicesecs:DescribeTaskDefinition*(ECS supports no resource scoping for this action)ecs:RegisterTaskDefinitionarn:…:task-definition/<agent family>:*ecs:UpdateServiceArnLike ecs:task-definition = arn:…:task-definition/<agent family>:*iam:PassRoleiam:PassedToService = ecs-tasks.amazonaws.comThe
ecs:task-definitioncondition stops the service from being pointed at another family's task definition, along with that family's roles. IAM Access Analyzervalidate-policyreturned no findings on the rendered policy.updaterrequiresruntime_updates. This is a variable validation, soupdater = trueon its own is refused at plan. Withoutruntime_updates, the next apply would put the module'simageback over the one the updater deployed.runtime_updatesreads the existing service, so both can only be enabled after the first apply. The README says so. The updater docs in the agent repo should say "setruntime_updates = trueandupdater = true" rather thanupdater = truealone.Verified
terraform testpasses 24 of 24: the existing 18 plus 6 new ones intests/updater.tftest.hcl. The new tests cover off by default, theruntime_updatesrequirement, channel validation, the full execution- and task-role policy JSON, PassRole limited to the two agent roles, the container's environment and secret, and the image defaults and override. They are in their own file so that their apply starts from empty state: run-leveloverride_resourcedoesn't reach resources that earlier applies inmodule.tftest.hclalready created.terraform fmt -check,terraform validate(module andexamples/basic),tflint,typosandprettier --checkall pass. terraform-docs v0.19.0 + prettier output is unchanged on a rerun. Terraform used locally: 1.12.2.Depends on, and not exercised before merge
elastio-database-monitoring-updateris published by the round-2 PR in elastio/database-monitoring-agent. Its ECR Public repos come from infraworld. The production default:0.1.10exists only onceagent-v0.1.10is released with it. Merge after that image is published.Please review; Anshuman will merge after approval.