Skip to content

Model Calibration - Partial Automation #717

Description

@joecastiglione

Develop software to automatically run the model, compare select sub-model results to expected values, and intelligently automatically adjust alternative specific constants based on the differences. For example, run trip mode choice, summarize trips by mode, compare trips by mode to the surveyed trips by mode, and adjust the mode specific constants based on the log of the difference ratio * a dampening factor. Repeat the procedure until difference is minimized and respect min and max adjustment constraints. This has been done before for simpler trip-based models and could be done for ABMs. This feature needs to be used with caution and would require design related to how to specify tracked and expected outputs for each sub-model.

May also want to consider Automated Calibration Pilot. Calibrating an activity-based model is challenging, as the number of parameters is large and the observed data is often scarce. In Version 1.6, the Consortium will begin experimenting with automated calibration procedures via one or more pilot demonstrations. The Consortium expects to continue improving the automated calibration procedures in subsequent ActivitySim releases.

Activity

  1. joecastiglione commented on Sep 15, 2023

    @joecastiglione
    Author

    Revised Description:

    The purpose of this work element is to develop either a stand-alone software tool or an integrated ActivitySim core capability, that provides the ability to automate the calibration of selected model coefficients and alternative specific constants. This capability must be able to run user selected model components and compare estimated results with the observed (or expected / asserted) values and mathematically adjust one or more parameters and/or alternative specific constants. This capability not intended to replace the value and central importance of more detailed examinations and comparisons between observed and estimated values. Rather, it is to provide a more efficient method for reaching the point where these more detailed comparisons can be made and a wide range of solutions considered where the model under development fails to achieve an acceptable level of calibration.

    This feature should have the ability to address the entire scope of ActivitySim model component types and component type elements. For example, the feature should be able to address basic calibration use cases where model system component calibration typically involves adjusting one or more alternative specific constants, such as for the auto ownership model, as well as more complex calibration use cases such as where calibration involves adjusting distance polynomial terms in a destination choice model. The features should provide clear, user-friendly summary statistics and displays, including comparative summaries of observed and estimated values (shares or absolute values) for each alternative specific constant in both tabular and graphic forms, by iteration, allowing the user to determine if the value of the constant displays a monotonic pattern or oscillates across the iterations

    Plan to design and implement for two different types of models: 1) an alternative specific constants model such as auto ownership and 2) a distances terms model such as work location choice.

    See MTC's specification for more detail.

  2. joecastiglione commented on Sep 15, 2023

    @joecastiglione
    Author

    Agency Comments:

    "Develop software to automatically run the model, compare select sub-model results to expected values, and intelligently automatically adjust alternative specific constants based on the differences. For example, run trip mode choice, summarize trips by mode, compare trips by mode to the surveyed trips by mode, and adjust the mode specific constants based on the log of the difference ratio * a dampening factor. Repeat the procedure until difference is minimized and respect min and max adjustment constraints. This has been done before for simpler trip-based models and could be done for ABMs. This feature needs to be used with caution and would require design related to how to specify tracked and expected outputs for each sub-model.

    May also want to consider Automated Calibration Pilot. Calibrating an activity-based model is challenging, as the number of parameters is large and the observed data is often scarce. In Version 1.6, the Consortium will begin experimenting with automated calibration procedures via one or more pilot demonstrations. The Consortium expects to continue improving the automated calibration procedures in subsequent ActivitySim releases. "

  3. moved this to Under consideration in Features for future developmenton Sep 25, 2023
  4. joecastiglione commented on Nov 23, 2025

    @joecastiglione
    Author

    From David Hensle (RSG):

    Model Calibration – Partial Automation

    Introduction
    Model calibration tools are extremely useful in the development and maintenance of travel demand models. This is especially the case in light of rapidly changing travel behavior and emerging modes of mobility. These factors have motivated many agencies to conduct regularly recurring cross-sectional travel surveys supplemented with passive data in order to monitor travel trends and perform more regular travel model updates. Activity-based model calibration is particularly complex given the number of model components and the interdependencies between them. Currently each agency (or their consultant) with an existing ActivitySim deployment has their own bespoke model calibration toolkit. For example, RSG has developed Jupyter notebooks for model calibration for several consortium members (SEMCOG, MWCOG, MetCouncil, SANDAG, TransLink, and CMAP), but the coverage of the notebooks is incomplete since some models did not require any calibration and the notebooks are not fully integrated with the ActivitySim software.

    Task 1: Design

    Considerable thought must be given to the development of partially automated calibration procedures. The procedures must consider:

    • Defining which alternative-specific constants to calibrate for each model. In some cases, multiple sets of constants may wish to be calibrated; for example, an auto ownership model could have constants for each auto alternative (except for the base) as well as 0-auto constants by district. A mode choice model may have constants by auto sufficiency and purpose, as well as a set of constants that consider premium transit usage.
    • Methods for calculating constants, such as the standard logit formulation (Adjustment = ln(obs/est)), a more complex form of constant adjustment (ln(((obsest)-obs)/((obsest)-est)), or the use of machine learning methods to calculate constant adjustments. The use of damping factors and methods to adjust damping factors in each iteration.
    • Minimum and maximum values of constants and other calibrated parameters.
    • The definition of convergence.
    • The use of observed data, including guidelines for the combination of household, on-board transit, and passive data for calibration. Data quantity will also be addressed – for example implications for calibration of alternatives and market segments with low numbers of observations and how that might affect the above issues.
    • Methods to check for reasonableness and present those checks to the user so that manual intervention can be applied during the calibration process.
    • Potential bundling of model components into calibration blocks that can be simultaneously calibrated to save time, such as a combination of tour frequency and stop frequency.
    • The software implementation including file and data management, performant calculations, reporting, and visualization.

    A major component to this work is developing a comprehensive design that addresses the above features. We anticipate multiple meetings with the consortium to lay out our ideas, provide examples of how it would work for certain sub-models, specify the input data requirements for model specification and potential variations of that input data, and what the output metrics and plots would be. These presentations would also include potential calibration settings such as which parameters will be adjusted during calibration, the allowable minimum, and maximum values of those parameters (possibly in relation to other parameters like in-vehicle time in mode choice), the treatment of unavailable alternatives (and very low frequency alternatives), the metrics for assessing goodness-of-fit and stopping criteria for calibration, and the user interface for the script(s).
    We see the following as key features of a successful calibration model in ActivitySim. First is the ability to iterate over one or more model components and adjust the coefficients. To properly adjust the coefficients, the model outputs would need to be summarized into a format that matches the targets and one that is consistent with how the coefficients are set up. Output summaries are needed at each iteration. These include, but are not necessarily limited to, a spreadsheet showing model counts, target counts, the coefficient adjustments, and the updated coefficient values, summary plots comparing model and target distributions, and a plot of coefficient adjustments over multiple iterations to show and test for convergence. Calibration functionality would need to be automated in the sense that the user should be able to launch the calibration and ActivitySim will be run iteratively, adjusting coefficients automatically for each iteration without the need for manual user intervention beyond the initial setup.

    To accomplish the above stated goals, we anticipate a design with roughly the following features. First, a calibration “configs” folder that will contain the calibration options. This includes a calibration.yaml that is an overall control module defining what model(s) to calibrate, the number of iterations and/or other stopping criteria, and other global options for calibration. We believe a useful feature would be to allow the user to specify multiple models and run them iteratively together or separately. For example, workplace location often plays a big role in tour mode choice and tour mode choice plays a big role in trip mode choice. After some iteration of each model individually, it would be desired to iterate all three models sequentially to converge on an overall calibration solution.

    Individual model calibration yaml files can also be included to specify model-specific settings. For example, in mode choice model calibration, it is often necessary to group together transit modes by access type. Another example would be calibration of work from home by resident county. The submodel yaml files will also point to the input target data file that should be used to calculate coefficient changes.

    Upon running ActivitySim in calibration mode, the output folder will contain a “calibration” subfolder that contains output for each iteration. This would include a spreadsheet that lists all calibration coefficients in the model, the target and model counts associated with that target, and the update applied, comparison plots between the model and the target counts, segmented according to the user options in the submodel yaml file, and an updated coefficients.csv file that will be used in the subsequent iteration. We anticipate building the iteration functionality within ActivitySim such that skims would not need to be re-loaded in each iteration, which would make performing iterations much faster.

    We expect there to be roughly two primary calibration summary procedures, one for ActivitySim’s “simple simulate” models and another for the “interaction simulate” models. Simple simulate models have alternatives specified as columns in the model specification. These include modes such as auto ownership and work from home. Interaction simulate models have many alternatives and the alternatives are not enumerated in the model specification document. These include models like destination choice, scheduling, school escorting, and the vehicle type model. We anticipate being able to build some very generalizable rules to summarize the simple simulate model outputs for calibration considering the alternatives are listed in the model specification. The interaction simulate models are much more likely to require code specific to the individual submodel, e.g. summarizing destination choice will look different than summarizing the school escorting model. As such, interaction simulate models will likely require more complex solutions.

    While most model components are relatively straightforward to calibrate, others (such as destination and mode choice) require more thought given that often calibration can occur in multiple dimensions. For example, while destination choice models are typically calibrated by distance, some models use piece-wise linear distance terms while others use non-linear terms which are typically linearized and regressed against to calculate calibration factors. District-level origin-destination flows may also be a target for destination choice calibration. Similar issues apply in mode choice. The design document will lay out agreed upon calibration framework for these models that can be applied across ActivitySim implementations.
    Each region often has its own unique calibration needs and may require their own specific solutions. To accommodate this, the design will consider a feature that allows the user to supply their own python code to summarize ActivitySim model outputs as they see fit without having to change the ActivitySim core code or calibration functionality. One way to implement this is to allow the user to point to a python file in the calibration yaml file that contains a summary function. ActivitySim will then use the custom summary function if supplied and default to the implemented example if not. This allows the user to have virtually unlimited flexibility and control over how to summarize their model for calibration.

    Deliverables:

    • Automated model calibration design document

    Estimated Level-of-Effort:

    • 60 hrs for the primary firm + 8 hrs each for other firms for in-depth review and discussion = ~92 hrs = ~$20k

    Task 2: Initial Implementation and Testing

    Calibration procedures will be implemented for three components of the chosen example model: auto ownership, workplace location choice, and tour mode choice. This provides coverage across simple simulate and interaction simulate models. The initial implementation will encompass building the framework for automated calibration integration into ActivitySim according to the design documentation and applying that framework to example submodels. Lessons learned from this initial implementation task can then be applied to future phases of ActivitySim development that will provide functionality to automatically calibrate each model component. The scripts will be implemented and tested using real-world survey data and implemented models at sufficient sample sizes to demonstrate that the calibration script works as expected.

    Deliverables:

    • Pull Request with updated code
    • Presentation to consortium showing calibration outputs

    Estimated Level-of-Effort:

    • ~80 hrs for ActivitySim calibration framework
    • ~80 hrs for implementation of 3 submodels
    • ~ 40 hrs for testing with real survey data
    • Total: ~200 hrs = ~$40k
  5. guyrousseau commented on Nov 24, 2025

    @guyrousseau

    So, for mode choice, a self-calibration subroutine will be added in the ActivitySim code to allow the user to enter calibration target values and ActivitySim will adjust coefficients until the maximum number of iterations is reached or the results fall within a user specified closure criteria. This feature will be added due to the amount of time calibration takes and the expectation that there would be numerous calibrations necessary as work progresses at an MPO during ActivitySim model development. In this case, the subroutine should also include summarizing the available person trips and compute an implied mode share using observed targets. This will help shed some light on the target values as well as potential issues with models that are upstream of mode choice. It will indicate certain market segments where mode choice has to compensate for inaccuracies occurring upstream.

  6. bhargavasana commented on Sep 8, 2026

    @bhargavasana
    Collaborator

    @joecastiglione this might be a good issue to edit and match to a template. This might potentially be closed by #1107.

  7. moved this from Under consideration to Funded in Features for future developmenton Sep 8, 2026
  8. joecastiglione commented on Sep 14, 2026

    @joecastiglione
    Author

    My sense is that #1107 substantially addresses this, but if not then yes, this could be an example issue to test against the template.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Feature (Phase)Feature for consideration in funded development phases

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions