Skip to content

database-monitoring: run the deployment updater on Fargate as a second ECS service (module 0.3.0) - #129

Merged
anchoo2kewl merged 2 commits into
masterfrom
feat/dbmon-updater-service
Oct 2, 2026
Merged

anchoo2kewl merged 2 commits into
masterfrom
feat/dbmon-updater-service

Conversation

@anchoo2kewl

@anchoo2kewl anchoo2kewl commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Closes #128.

Fargate deployments have no host to run the deployment updater on, so "Update to latest" never appears for them. With updater = true the module now runs the updater as a second ECS service in the agent's cluster.

New inputs: updater (default false), updater_image (default by channel: public.ecr.aws/elastio/elastio-database-monitoring-updater:0.1.10 for production, public.ecr.aws/elastio-development/elastio-database-monitoring-updater:latest for development), update_channel (production | development), agent_id (default "", required and validated as a UUID when updater = true: the updater cannot discover it, and without it the service would run and never update anything; the agent logs it on every start as registered as <id>). New outputs: updater_service_name and updater_task_role_arn. The module version goes to 0.3.0. That release also carries runtime_updates from #127, which merged after 0.2.1 was tagged.

What updater = true creates: one service, <name>-updater (desired 1, minimum healthy 0 / maximum 100 so two never run at once), with a 0.25 vCPU / 512 MiB ARM64 task. It runs in the agent's subnets and security groups and needs no inbound rules. It logs to the module's log group under updater/. Its environment is ELASTIO_DBMON_SERVER_URL, _AGENT_ID, _DEPLOYMENT=ecs, _UPDATE_CHANNEL, _ECS_CLUSTER, _ECS_SERVICE and _ECS_CONTAINER, plus ELASTIO_DBMON_API_KEY from the agent's own secret.

The service also gets its own two roles:

  • Execution role: AmazonECSTaskExecutionRolePolicy, the same as the agent's (image pull and log write), plus secretsmanager:GetSecretValue on the agent's API-key secret only.
  • Task role, the whole policy:
Sid Action Resource Condition
ReadAgentService ecs:DescribeServices the agent's service ARN
ReadTaskDefinitions ecs:DescribeTaskDefinition * (ECS supports no resource scoping for this action)
RegisterAgentTaskDefinition ecs:RegisterTaskDefinition arn:…:task-definition/<agent family>:*
DeployAgentService ecs:UpdateService the agent's service ARN ArnLike ecs:task-definition = arn:…:task-definition/<agent family>:*
PassAgentRoles iam:PassRole the agent's task role and execution role iam:PassedToService = ecs-tasks.amazonaws.com

The ecs:task-definition condition stops the service from being pointed at another family's task definition, along with that family's roles. IAM Access Analyzer validate-policy returned no findings on the rendered policy.

updater requires runtime_updates. This is a variable validation, so updater = true on its own is refused at plan. Without runtime_updates, the next apply would put the module's image back over the one the updater deployed. runtime_updates reads the existing service, so both can only be enabled after the first apply. The README says so. The updater docs in the agent repo should say "set runtime_updates = true and updater = true" rather than updater = true alone.

Verified

  • terraform test passes 24 of 24: the existing 18 plus 6 new ones in tests/updater.tftest.hcl. The new tests cover off by default, the runtime_updates requirement, channel validation, the full execution- and task-role policy JSON, PassRole limited to the two agent roles, the container's environment and secret, and the image defaults and override. They are in their own file so that their apply starts from empty state: run-level override_resource doesn't reach resources that earlier applies in module.tftest.hcl already created.
  • Each of these seeded violations fails a test: a third PassRole ARN, dropping the UpdateService condition, and neutering the validation.
  • terraform fmt -check, terraform validate (module and examples/basic), tflint, typos and prettier --check all pass. terraform-docs v0.19.0 + prettier output is unchanged on a rerun. Terraform used locally: 1.12.2.

Depends on, and not exercised before merge

  • The updater image. elastio-database-monitoring-updater is published by the round-2 PR in elastio/database-monitoring-agent. Its ECR Public repos come from infraworld. The production default :0.1.10 exists only once agent-v0.1.10 is released with it. Merge after that image is published.
  • No real apply has been run. The IAM policy was checked by Access Analyzer and the mocked tests, not against a live ECS UpdateService call.

Please review; Anshuman will merge after approval.

…d ECS service (module 0.3.0)

New inputs updater, updater_image, update_channel and agent_id. With
updater = true the module runs elastio-database-monitoring-updater as a
second service in the agent's cluster, with its own execution role (reads
only the agent's API-key secret) and a task role that may only describe and
update the agent's service (to a revision of the agent's family), register
that family, describe task definitions, and pass the agent's two roles to
ECS. updater requires runtime_updates.

Closes #128
Copilot AI balanced review requested due to automatic review settings October 2, 2026 18:32

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@maksvet
maksvet self-requested a review October 2, 2026 19:19
variable "agent_id" {
description = "The agent's ID in Elastio, given to the updater. Optional: when empty, the updater asks Elastio for it with the agent's API key."
type = string
default = ""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

agent_id must be required when updater = true.
The updater does not ask Elastio for the id. In update-agent.py, an empty ELASTIO_DBMON_AGENT_ID only logs "Updates are off until then" every 5 min.
So with the default "" the updater service runs and costs money, but never updates the agent.
Please add a validation like the runtime_updates one, and fix the agent_id description, the README and the PR body.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 41b2b2b: agent_id is now validated like runtime_updates — updater = true without a UUID agent_id is refused at plan (a non-UUID too; it stays optional with the updater off). Description, README (inputs table regenerated, example and prose) and the PR body are corrected; they now point at where the ID is found: the agent logs registered as <id> on every start in the module's log group. Three new tests (missing, not a UUID, optional when off); 27/27 pass, and neutering the validation fails updater_requires_agent_id.

The updater cannot discover the agent's ID: with an empty one it runs,
costs money and never updates the agent (review on #129). Validate it like
runtime_updates, point the description and README at where the ID is
found (the agent logs "registered as <id>" on every start), and test the
refusal, a non-UUID, and that it stays optional with the updater off.

@maksvet maksvet left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. The agent_id fix is good.
Not blocking: the new .terraform.lock.hcl at the module root can go. In this repo only examples/ keep lock files.

@anchoo2kewl
anchoo2kewl merged commit 864e94d into master Oct 2, 2026
22 checks passed
@anchoo2kewl
anchoo2kewl deleted the feat/dbmon-updater-service branch October 2, 2026 21:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

database-monitoring: run the deployment updater on Fargate as a second ECS service

3 participants