2  Lab Notebook

This is a running record of meeting notes made for the execution of this project

2.1 To Do

  • finalize prism documentation

2.2 Notes

EMA workflow:
1. Original plan was for PRISM to update ema in real time as participants loaded the survey. This requires the use of a reverse proxy (ngrok) which will not be approved for use by Cybersecurity 2. Revised plan was to have each participant have a single link for ema, which is updated daily by the API. Discovered 2026_0409 that although an update endpoint exists, it can ONLY update a very few specific fields and cannot update the whole survey. 3. New plan: Use a qualtrics contact list with custom fields. Each participant will have a unique link to the ema survey. API can update the custom fields with lat/lon for uncontextualized places. We should probably save a version of this file dated each day so that we have a backup record of exactly who got what. a. Download ema and add context to places file. b. Update contact list with new uncontextualized places. c. Push up. d. Need to modify PRISM to pull Twilio links from this file.

2.3 Lab Meeting Notes

The following are executive summaries of transcripts and notes, summarizing decision points and tasks, provided by copilot. Full notes, transcripts, and recordings are available in S:optimize/administration/lab_meetings

Note: The procedure for obtaining transcripts and recordings is

  • Recording the meeting with the Android Recorder App automatically provides a .txt transcript.
  • If you record with other software and have an mp3/m4a audio file, use the wisc.edu Office 365 instance of Word. Open a blank document and you will see a “Dictate” item on the toolbar. Click that and choose Transcribe, and it allows you to upload the audio file and then provides a transcript. When finished, hit “Add to document -> include both speakers and timestamps” to place the transcript in the blank document, and then save it (name it the date of the transcript).
    • NB copilot will cheerfully tell you that it can transcribe an uploaded audio file, but that is incorrect: audio file types are not a permitted upload type.
  • Login to the wisc.edu instance of copilot with your netID. Upload the transcript (you can select the cloud document that you just saved) and ask copilot: “please provide a very detailed summary of the attached meeting transcript in markdown format”.
  • When it provides the summary, click the copy icon at the bottom and it will copy it in markdown format for pasting into a qmd. You’ll have to do a small bit of cleanup (removing headers, section breaks, and footers) to paste here.

2.3.1 2026_0430

EMA Emotion Rating Discussion

A significant portion of the meeting focused on how to explain EMA emotion questions to participants.

  1. Purpose of the EMA Emotion Measure

The EMA (Ecological Momentary Assessment) emotion questions are intended to capture participants’ momentary affective state in a way that is:

  • Simple enough for repeated daily use
  • Theoretically grounded in affect science
  • Compatible with modeling approaches
  • Minimally cognitively demanding for participants

The discussion emphasized that EMA responses are not meant to capture discrete emotion labels (e.g., “fear,” “sadness”) but rather a continuous affective state that can be quantified and modeled over time.

  1. Two-Dimensional Affect Model Underlying the EMA

2.1. Core Dimensions

The EMA emotion measure is explicitly grounded in a two-dimensional model of affect, consisting of:

  1. Pleasantness (Valence)
    • Ranges from unpleasant / unhappy to pleasant / happy
  2. Activation (Energy/Arousal)
    • Ranges from still / quiet / low energy to energized / activated / high energy

Participants answer two separate questions, one for each dimension, rather than a single combined question.

2.2. Independence of Dimensions (Critical Concept)

A key conceptual point repeatedly emphasized is that pleasantness and activation are independent dimensions:

  • A feeling can be:
    • Pleasant and high activation (e.g., excitement, joy)
    • Pleasant and low activation (e.g., calm, contentment)
    • Unpleasant and high activation (e.g., fear, alarm)
    • Unpleasant and low activation (e.g., depression, fatigue)

This independence is central to:

  • Accurate emotional modeling
  • Avoiding forced assumptions about “intensity”
  • Allowing nuanced emotional states to be represented.
  1. Challenges in Explaining the EMA to Participants

3.1 Cognitive Load and Over-Explanation Risk

The group identified a strong risk that:

  • Overly technical explanations (e.g., circumplex models, rotated axes)
  • Dense visuals
  • Excessive terminology

would overwhelm participants, particularly in an EMA context where questions must be quick and intuitive.

A repeated concern was that what makes sense to researchers does not necessarily translate to participants.

3.2. Problems with Emotion Labels and Discrete Categories

Several issues were identified with using discrete emotion words too heavily:

  • Participants may feel conflicted if they experience multiple emotions simultaneously (e.g., tired and frustrated).
  • Discrete labels imply exclusivity or dominance that does not reflect lived experience.
  • Emotion words can differ in meaning across individuals.

As a result, the group emphasized:

  • EMA questions should not require participants to name or choose a single emotion
  • Emotion words should be used only illustratively, not as response options.
  1. Visual Design Decisions: What Not to Show

4.1. The Circumplex Model

Although the circumplex model of emotion is theoretically accurate and useful for researchers, the group agreed that:

  • Participants do not need to see a circle
  • Explaining why emotions form a circle introduces unnecessary complexity
  • Boundaries and geometry invite confusion (“Why a circle?” “Why these limits?”)

Decision: Susan will re-format to not show a circumplex or circular model to participants.

4.2 Two-Axis Grids

Similarly, displaying a full 2D grid with intersecting axes was judged problematic:

  • Participants will never see a grid during actual EMA completion
  • The grid introduces a translation step (grid → sliders) that adds confusion
  • Numeric tick marks and coordinates are meaningless in the EMA interface

Decision: Avoid grids and focus instead on the actual sliders participants will use.

  1. Recommended EMA Explanation Strategy

5.1. Anchor the Explanation in the Actual Questions

The group strongly recommended aligning training materials with the exact EMA interface:

  • Show the two slider questions directly
  • Emphasize that they are separate, independent ratings
  • Avoid abstract terminology unless absolutely necessary

Suggested framing:

  • “You are going to answer two questions about how you feel right now.”
  • “These two questions describe different aspects of how you’re feeling.”.

5.2. Using Example Emotions Carefully

Emotion words should be used only as concrete examples to illustrate how the dimensions work:

  • Example contrasts discussed:
    • Fearful vs. depressed
      • Both unpleasant
      • Fearful = high activation
      • Depressed = low activation
    • Overjoyed vs. tranquil
      • Both pleasant
      • Overjoyed = high activation
      • Tranquil = low activation

The purpose of these examples is not classification, but to demonstrate that:

  • Pleasantness ≠ energy
  • Unpleasantness ≠ intensity.

5.3. Avoiding “Intensity” Language

The group identified “intensity” as potentially misleading:

  • High activation is not the same as “stronger emotion”
  • Low activation emotions can still be highly important or impactful

Recommended language:

  • Use energy, activation, or how still vs. energized you feel
  • Avoid asking participants to focus on “the most intense emotion”

Instead, participants are encouraged to:

  • Report how they feel overall in the moment
  • Provide a subjective snapshot, not an emotional diagnosis.
  1. Handling Mixed or Simultaneous Emotions

A recurring concern was how participants should respond if they feel multiple emotions at once.

Key clarifications:

  • EMA ratings are not averages over named emotions
  • Participants do not need to resolve emotional conflicts explicitly
  • They are simply rating their current experiential state along two dimensions

The group described this as:

“Where you land in this space right now,”
not
“Which emotion is dominant.”

This framing reduces anxiety about “getting it right”.

  1. Instructional Tone and Participant Experience

7.1. Emphasizing Subjectivity

Participants should be clearly told:

  • There is no correct answer
  • People may experience and rate emotions differently
  • They should answer based on their own experience, not how emotions “should” be rated

This was seen as critical for:

  • Data quality
  • Participant confidence
  • Reducing response fatigue.

7.2. Live or Interactive Examples (Optional)

The group discussed—but did not commit to—interactive training elements such as:

  • Asking participants to think of a recent situation
  • Walking through how they might rate it on both sliders

While potentially helpful, the consensus leaned toward:

  • Keeping the explanation lightweight
  • Avoiding excessive pre-training that could discourage participation.
  1. Final Conceptual Takeaway

The EMA emotion measure is designed to:

  • Capture momentary affective state, not emotion labels
  • Use two independent dimensions
  • Minimize cognitive burden
  • Align tightly with the participant-facing interface

The guiding principle is:

“Help participants understand just enough to answer confidently, without turning the explanation itself into a task.”.

2.3.2 2026_0326

Key Decisions

  1. Feature Grouping Strategy
    • Retain the expanded feature set (≈51 features) grouped into conceptual categories.
    • Avoid splitting past day vs. past week features in most cases, as the distinction adds limited clinical value.
    • Maintain a max vs. median distinction only for pleasant events, since they have opposite risk directions.
  2. Emotion-Based Location Features
    • Acknowledge that pleasant and unpleasant locations appear protective, while mixed-emotion locations increase risk.
    • Treat this pattern as interpretable (re‑engagement, preparedness, predictability), not a modeling error.
    • Likely move toward separating pleasant, unpleasant, and mixed-emotion locations into distinct categories for more precise interventions.
  3. Efficacy/Motivation Measures
    • Combine most recent and past‑week efficacy/motivation EMA measures into a single category rather than keeping them separate.
    • Continue to treat abstinence commitment as a standalone protective factor.
  4. EMA Design Philosophy
    • Simplify daily EMAs to reduce burden and confusion.
    • Prioritize actionable, interpretable signals over granular but noisy inputs.
  5. Model Performance Interpretation
    • Accept that GPS features are now appearing more frequently as important predictors after recent fixes.
    • Focus on representational balance (variety of feedback) rather than maximizing predictive gain alone.

Action Items

  1. Emotional Location Analysis (Claire)
    • Split location EDA tables by pleasant, unpleasant, mixed, neutral in side‑by‑side columns.
    • Identify where unpleasant locations meaningfully diverge from overall trends.
    • Share updated comparison tables with the team.
  2. Feature Reconfiguration (John)
    • Separate emotional-location categories (pleasant / unpleasant / mixed) in the model and re‑evaluate their frequency and impact.
    • Remove or down‑weight very small features (e.g., past‑day work exposure).
    • Recombine efficacy/motivation day vs. week measures into a single category.
  3. EMA Interface Updates (Susan)
    • Redesign drinking report:
      • Ask only for hour drinking started (not duration).
      • Fix standard drink scale (include missing category).
    • Replace or prototype alternatives to sliders (e.g., Likert buttons).
    • Test usability of revised EMAs on mobile devices.
  4. Baseline Measure Review (Susan → John)
    • Compile baseline measures from Risk‑1 (e.g., alcohol history, self‑efficacy, craving, stress, social support).
    • Flag items that could help impute missing new EMA variables (motivation, confidence, social support).
  5. Imputation Strategy (John)
    • Develop imputation models for new EMA items using:
      • Existing EMA correlations (craving, stress, efficacy).
      • Selected baseline measures.
    • Avoid default median imputation when person‑specific predictors are available.
  6. Location Survey Compatibility
    • Re‑add legacy GPS questions where needed:
      • “Is this a risky place?”
      • “Are you trying to avoid this place?”
    • Keep new harm/help and emotional grid questions in parallel to enable mapping between old and new features.
  7. Follow‑Up Preparation
    • Circulate revised tables, EMA mockups, and baseline item lists via Slack.
    • Schedule follow‑up discussion focused on:
      • Final emotional-location handling.
      • EMA usability feedback.
      • Readiness for training next‑round models.

Detailed Notes

1. Overview of Model Results and Feature Retention

1.1 The team reviewed a predictive lapse-risk model retaining 51 features, most of which are intuitive and interpretable.

1.2 Features include GPS-derived location behaviors, EMA self-reports, temporal patterns, and emotional/psychological indicators.

1.3 These features were grouped into 28 conceptual categories, though participants discussed whether some categories should be further split or collapsed.

2. GPS-Based Location Features and Their Associations

2.1 Time spent in locations where alcohol is present (past day and past week) is a clear risk factor.
2.2 Time spent in places the participant wants to avoid (past week) is also a risk factor.
2.3 Time spent in locations where the participant drank in the past (past day and week) increases risk.
2.4 High-risk and medium-risk locations increase risk, especially over the past day and week.
2.5 AA meetings, friends’ homes, healthcare settings, and being at home are protective.
2.6 Restaurants show a surprising pattern: time spent there (past day/week) is associated with increased risk.
2.7 Work locations show very small effects, with past-week work exposure slightly increasing risk and past-day exposure slightly decreasing risk; discussion suggested this may simply reflect day-of-week effects rather than meaningful causal relationships.
2.8 Location entropy/variance (irregular schedules, being “all over the place”) is consistently risky.

3. Emotional Location Ratings: Pleasant, Unpleasant, Mixed

3.1 Locations were labeled by participants as pleasant, unpleasant, mixed, or neutral (neutral dropped from the model).
3.2 Mixed-emotion locations modestly increase risk.
3.3 Pleasant and unpleasant locations are protective, with unpleasant being particularly counterintuitive.
3.4 The team flagged unpleasant locations as the least intuitive finding, deserving further investigation.
3.5 Exploratory analysis showed unpleasant locations are disproportionately associated with high-risk categories (e.g., liquor stores/bars) but may also reflect re-engagement with real-life obligations (errands, work).
3.6 Hypotheses for unpleasant locations being protective included:

  • Preparedness/vigilance when entering known unpleasant environments
  • Reconnecting with avoided but meaningful life domains (family, responsibilities)
  • Reduced emotional unpredictability compared to mixed-valence locations

3.7 The group debated whether pleasant, unpleasant, and mixed should remain combined or become separate intervention categories, with growing support for splitting them.

4. EMA-Based Risk, Stress, and Motivation Features

4.1 Craving

  • Maximum craving over the past week is a strong risk factor.

4.2 Risky Situations

  • Both most recent and maximum risky situations are associated with higher lapse risk.
  • Past and future risky situations were separated into different conceptual categories to support tailored feedback.

4.3 Stressors/Hassles

  • Maximum stress over the past week and most recent stressor reports increase risk.
  • Future anticipated stressors are also risky.

4.4 Pleasant Events

  • Median pleasant event ratings over the past week are protective.
  • Maximum pleasant event ratings are risky, likely reflecting celebrations or partying.
  • This distinction was considered clinically actionable, prompting either:
    • Increasing regular, healthy pleasant activities (low median), or
    • Reflecting on extreme celebratory events (high max).

4.5 Negative Affect

  • Activated negative affect (anger, irritation) increases risk.
  • Deactivated negative affect (sad, depressed) is an even stronger risk factor.
  • Low activation positive affect is interpreted as high risk via the opposite pole of the affect continuum.

4.6 Past Alcohol Use (Lapses)

  • Any recent drinking or lapse history strongly elevates current risk.

5. Efficacy, Motivation, and Abstinence Commitment

5.1 EMA item predicting likelihood of drinking next week shows strong predictive power.
5.2 Commitment to abstinence (coded no / uncertain / yes) is protective.
5.3 The team debated splitting most recent vs. past-week averages for efficacy/motivation, ultimately leaning toward combining them to reduce complexity and redundancy.

6. Time and Calendar Effects

6.1 Friday and Saturday are associated with increased risk.
6.2 Monday and Tuesday are protective.
6.3 These effects likely capture social norms and routine structure rather than individual behavior alone.

7. Category Selection, Frequency, and Personalization

7.1 A “max feature per day” approach with secondary-category selection was reviewed.
7.2 A second category is selected ~21% of days under the 25th-percentile importance criterion.
7.3 Most frequent top categories include:

  • Efficacy/motivation
  • Craving
  • Future risky situations
  • Past use
  • Medium-risk locations
  • Abstinence commitment 7.4 Rare categories include work, healthcare, and friends’ homes.
    7.5 Across participants:
  • Minimum categories seen over 3 months: ~3
  • Maximum: ~16
  • Average: ~10, considered reasonable and comparable to structured therapy manuals. 7.6 Long consecutive runs of the same category occur mainly in participants with very few total categories.

8. Decisions About Day vs. Week Distinctions

8.1 The group largely agreed not to distinguish between past-day and past-week features, except where:

  • Directionality differs (e.g., pleasant event median vs. max)
  • Clinical interpretation clearly changes
    8.2 Day vs. week distinctions may add noise rather than actionable insight.

9. Improvements to EMA Design (Daily Surveys)

9.1 Drinking-report interface issues were discussed:

  • Hour-by-hour scrolling was confusing (especially post-midnight drinking).
  • Decision made to simplify by asking only the hour drinking started, not duration. 9.2 Standard drink count scale needed correction (missing category). 9.3 Sliders were criticized as:
  • Annoying on mobile devices
  • Prone to accidental or default responses 9.4 Alternatives proposed:
  • Likert-style buttons with no default selection
  • Clearly labeled anchors with mandatory response 9.5 Emotional arousal/valence (circumplex model) will be retained but with improved participant orientation.

10. New EMA Questions and Imputation Challenges

10.1 New EMA items added:

  • Social support
  • Future pleasant events
  • Separate efficacy, motivation, and confidence questions 10.2 These new items are missing in older datasets (Risk-1). 10.3 Without imputation, they would be weak predictors due to missingness.
    10.4 Proposed solutions:
  • Impute using other EMA responses (e.g., craving, stress)
  • Use baseline measures correlated with motivation/confidence

11. Baseline Measures Reviewed for Imputation Potential

11.1 Candidate measures included:

  • Alcohol use history and past quit attempts
  • Alcohol Abstinence Self-Efficacy Scale
  • Craving scales
  • Depression and perceived stress scales
  • Quality of life and social support measures 11.2 DSM and family functioning measures were considered less directly useful. 11.3 Even weak-person-specific predictors were preferred over global median imputation.

12. GPS Location Survey and Feature Compatibility

12.1 The updated location context survey captures:

  • Location type
  • Alcohol availability
  • Helpfulness or harmfulness to recovery
  • Emotional valence via a 2D grid 12.2 Challenges identified:
  • Old features (“riskiness,” “trying to avoid”) are not fully captured.
  • Backward compatibility is needed to use past models. 12.3 Proposed solution:
  • Add legacy questions back in alongside new ones temporarily.
  • Use overlap data to map old features onto new formats.

13. Outstanding Questions and Next Steps

13.1 Split emotional categories (pleasant, unpleasant, mixed) for analysis.
13.2 Revisit the interpretation of unpleasant locations if they persist as protective.
13.3 Finalize EMA UI changes and test usability on phones.
13.4 Identify baseline variables most useful for imputing new EMA items.
13.5 Maintain consistency between GPS and EMA feature definitions across studies.

2.3.3 2026_0312

Meeting Focus Areas:

  • Discussion of a recent machine‑learning study on predicting alcohol lapses
  • Comparison to the group’s ongoing work (RISK, OPTIMIZE)
  • Feature‑selection critique
  • Fairness, subgroup modeling, interactions
  • EMA redesign for OPTIMIZE
  • Additional methodological considerations
  1. Overview of the Paper Being Discussed

The meeting starts with a discussion of a research paper analyzing heavy‑drinking prediction using machine‑learning models relying primarily on baseline characteristics rather than proximal/dynamic factors.

Key Observations:

  • Their reliance on baseline variables (static, individual‑difference measures) is conceptually different from the group’s approach, which emphasizes proximal, time‑varying predictors from EMA (e.g., craving, stress, context).
  • The paper reports substantial predictive signal from baseline factors, but the group stresses that predicting lapse timing requires proximal information they did not include.
  • The study also uses sex‑stratified models—building separate prediction models for participants identified male vs. female at birth—and finds differing sets of influential predictors.
  1. Discussion of Modeling Workflow in the Paper

2.1 Algorithms Used

  • Primary models appear to be Random Forests, possibly also tree‑based gradient boosters.
  • The group compares these to their standard models (Random Forest, XGBoost/ExtraBoost).

2.2 Feature Engineering

  • Baseline characteristics only—no time‑varying covariates.
  • Includes psychiatric measures, drinking history, social network characteristics, labs, demographics.

2.3 Feature Selection: Boruta Algorithm

A major discussion point.

Boruta uses random “shadow” features, permutes values, and retains real variables that repeatedly outperform these noise variables.

Group critique:

  • Tree‑based models already handle irrelevant features reasonably well.
  • Boruta may add variance since feature selection itself is a noisy process.
  • Interpretation benefits are unclear; the group argues that permutation importance on the full model is often more transparent.
  • They note the authors never provided justification for why feature selection was needed.
  1. Concerns About Imputation and Missing Outcomes

The paper treated all missing drinking observations as heavy drinking.

Group reactions:

  • This resembles “worst‑case” intent‑to‑treat assumptions in clinical trials, but is odd in this context.
  • The approach likely decreases predictive performance due to label noise.
  • Alternatives discussed:
    • Last‑observation‑carried‑forward
    • Dropping uncertain observations (as their own group typically does)
    • Conservative vs. practical trade‑offs
  1. Feature Differences by Sex at Birth

The paper finds different feature sets in male vs. female models.

4.1 For male participants:

Top features skew toward alcohol‑dependence severity, drinking history, and biomarkers (e.g., liver enzymes).

4.2 For female participants:

Top features include:

  • Psychiatric symptom measures
  • Emotional/mental health indicators

This difference leads to:

  • Discussion of interactions,
  • Potential fairness implications,
  • Importance of examining subgroup‑specific importance in the group’s own models.

However, the group warns:

  • Different features across stratified models may result from variance, sample‑size imbalance, or correlated predictors, not true moderation.
  1. Fairness & Interaction Effects in Prediction Models

5.1 Current status in the group’s existing models:

  • Prior RISK1 models showed fairness problems (e.g., poorer performance for women or certain racial groups).
  • RISK2 appears more fair but still shows ~0.02 AUROC advantage for men.
  • Demographic variables typically rank low in feature importance, suggesting little evidence of strong interactions.

5.2 Challenges Identified:

  • Tree‑based models can capture interactions, but:
    • They may miss small‑effect interactions.
    • Detection tools (e.g., Friedman’s H‑statistic) are underused or not well understood.
  • Fitting separate models per subgroup is conceptually appealing but methodologically messy.

5.3 Next Steps Proposed:

  • Examine feature importance within subgroups more systematically.
  • Investigate methods for identifying interactions in ML models (Friedman’s H, SHAP interaction values).
  • Explore impacts of subgroup representation (e.g., racial/ethnic imbalance).
  1. Discussion of Social Network Measures

Interest in incorporating social network variables (e.g., “percentage of drinkers in network,” “important person drinking status”), which showed importance in the paper.

Concerns:

  • Social network characteristics change over time, especially during early recovery.
  • Baseline‑only measures may have limited value in day‑level prediction.

Outcome:

  • Group agrees they should learn how these social metrics were measured and consider adding or adapting them in future studies.
  1. Relevance for OPTIMIZE: Incorporating Baseline Features

The group debates whether adding baseline psychiatric or dependence variables would meaningfully improve their own predictions.

Conclusions:

  • Baseline variables usually don’t improve AUROC.
  • They might, however, improve:
    • Calibration,
    • Spread of predicted probabilities,
    • Subgroup fairness.
  • Theoretical models (e.g., relapse prevention) suggest distal factors may shape proximal factors (via mediation rather than direct moderation).
  1. Redesigning the OPTIMIZE EMA

A major portion of the meeting discusses revising the daily EMA instrument.

2.3.4 8.1 Required Additions for Lapse Reports

New lapse‑related data to collect:

  • Day & time of drinking onset
  • Number of standard drinks consumed
  • Maintain existing questions: “Did you drink?” + “Do you still want to remain abstinent?”
  • Add drinking details per lapse event.

Handling multiple drinking episodes per day:

  • Group agrees:
    • Prefer capturing onset time(s) and total drinks.
    • Optional: allow multiple episodes (but complexity increases).
  • Ultimately:
    Collect at least onset time + drinks per episode, with flexibility to refine later.

8.2 Limiting retrospective window

Implementation constraints in Qualtrics:

  • Easier to allow reporting lapses only up to ~15 days back (not 30+).
  • Group feels lapses older than 2 weeks are unlikely to be reported reliably anyway.
  • 15 days is acceptable.
  1. EMA Content Review & Revisions

9.1 Keep as is:

  • Drinking report
  • Goals (abstinence)
  • Craving intensity
  • Risky situations
  • Hassles / stressors
  • Positive events
  • Mood (valence & arousal)
  • Motivation & confidence

9.2 Mood / Arousal Scale Revision

Current arousal anchors (“aroused/alert”) are misleading.

Proposed alternatives:

  • High energy ↔︎ low energy
  • Strong ↔︎ weak emotions
  • Provide examples across valence (e.g., excited vs. bored)

Team wants to:

  • Review published affect circumplex measures,
  • Possibly adopt a two‑axis mood grid,
  • Ensure lay‑friendly wording but good construct validity.

9.3 Additions Under Consideration:

  • Daily social support (“To what extent did you feel supported by friends/family today?”)
  • Clarifying stressful event examples (include interpersonal conflict)
  • Remove or avoid:
    • Sleep (has not emerged as important)
    • Daily companion drug‑use reports (rare, low variance)

2.4 10. EMA Length Considerations

  • New EMA will have ~16 questions (previous Risk‑2 EMA had ~10–12).
  • Once‑daily administration makes this acceptable.
  • Further optimization may occur once pilot data are reviewed.
  1. Next Steps / Action Items For Susan:
  • Implement lapse‑report section with:
    • Date
    • Onset time
    • Drinks consumed
    • Max 15‑day retrospective window
  • Add placeholder for social‑support item.
  • Begin revising arousal scale + mood anchors.
  • Show prototype of lapse‑reporting workflow.

For the team:

  • Review example implementations in Slack.
  • Investigate interactions in ML models for future fairness analyses.
  • Explore measurement options for social network metrics.

2.4.1 2026_0305

1. Meeting Context and Attendees

The meeting reconvened after last week’s discussion to review updates on the predictive modeling work for the alcohol‑use research project.
Attendees included:

  • Chris (leading the modeling work)
  • Kendra, Claire, Coco, Susan, Andrew (referenced), and others on the research team

The first few minutes included informal coordination about schedules, practica, and waiting for Claire to arrive from a clinic intake.

2. Modeling Updates (Version 26 → 28)

Chris presented extensive progress on model versions v26 through v28, primarily concerning how emotion variables (PA/NA), GPS data, EMAs, and risk features should be included in the lapse‑risk prediction model.

2.1 Switching from Quadratic PANAS to Linear PANAS

Previously, the model used quadratic terms for PA (Positive Affect) and NA (Negative Affect).
A problem was identified:

  • Quadratic PA/NA terms produce U‑shaped relationships — e.g., “risk is high if you’re either very sad or very elated.”
  • This is not easily interpretable when giving feedback to participants.
  • It leads to nonsensical participant messaging such as: > “Your risk is high because you’re either angry or content.”

Decision:
Move to linear PA/NA terms instead of quadratic ones because:

  • They are more interpretable.
  • Predictive accuracy stayed nearly the same (quadratic AUC ~0.863 vs. linear ~0.862).
  • The team prefers PANAS over valence/arousal due to interpretability and intervention relevance.

2.2 Reintroducing “Min” Features

“Min” values (minimum EMA levels across past week) were previously removed but now will be reintroduced because:

  • Low PA / high NA minima matter (e.g., depressive affect has predictive value).
  • Other features like “time at home” may also show important minimum-based effects.

Kendra flagged this, and Chris agreed to re‑add minimums after Claire finishes current tests.

2.3 GPS Feature Corrections

Several GPS‑related improvements were made:

2.3.1 Handling Missing GPS Correctly

Originally, if no GPS points existed for “locations where alcohol was present/drank,” the model incorrectly assigned 0 time at that location.

Issue:

  • 0 ≠ no risk.
    Often the GPS simply failed.

Fix:

  • Replace missing values with NA rather than 0 so they are correctly imputed (e.g., median imputation) rather than treated as 0 risk.

2.3.2 Potential Future Idea: Missingness as Feature

Kendra asked whether missing GPS might be meaningful.
Chris explained:

  • Missing EMA is meaningful (participant avoidance),
    but
  • Missing GPS is usually technical, not behavioral.
    Thus, missingness is not used as a predictive feature.

2.4 Transition to Total Effects Rather Than Direct Effects

Chris emphasized a major conceptual shift:

Direct GLM coefficients are misleading

Why?
Because in a logistic model, predictors can influence risk indirectly through other variables.
Example:

  • Stress → expected increase in drinking likelihood
    But holding “predicted drinking” constant can mask the real causal effect.

New approach

Calculate total effects by:

  1. Using model‑predicted probabilities for each observation
  2. Convert to odds → log‑odds
  3. Compute correlations between log‑odds and each predictor

Reason:

  • Log‑odds linearize the relationship, making them comparable.
  • Probability is nonlinear and misleading for reasoning about feature effects.

Outcome:

  • Feature directions now align neatly with domain expertise.
  • For example, stress becomes positively associated with lapse risk once total effects are used.

3. Features Retained in the V28 Model

The final v28 model includes:

Positive Predictors of Lapse Risk

  • Time at places where participant drank
  • Time at places where alcohol is present
  • Time in restaurants
  • Time in risky locations
  • Future predicted drinking (EMA 10)
  • Craving (EMA)
  • Stress (EMA)
  • Future risky situations (EMA 8)
  • Activated NA (anger, agitation)
  • Day‑of‑week effects (Friday, Saturday)

Negative Predictors of Lapse Risk

  • Time at home
  • Time at pleasant locations
  • Activated PA (energy, enthusiasm)
  • Monday effect
  • Abstinence goal commitment

The model is now consistent, interpretable, and aligned with clinical expectations.

4. Discussion: Personalized Messaging Logic

A major focus of the meeting was how to choose which feature to show participants each day.

Two strategies were compared:

4.1 Strategy 1: Always Use the Top Feature

  • Each day, the feature with highest log‑odds is chosen.

Pros:

  • Simple, interpretable

Cons:

  • Leads to long “runs” where the same feature is shown repeatedly (e.g., 40–70 days).
  • Not engaging for participants.

Many participants received only 2–4 distinct features over months.
Some had the same feature 50–70 days in a row.

This raised concerns about:

  • Engagement
  • Perceived personalization
  • Perceived accuracy of the system

4.2 Strategy 2: Allow Switching to the Second Feature

Rule:

  • If the same feature appears two+ days in a row,
  • AND the second‑most predictive feature is not too much worse,
    then switch.

Threshold used:

  • Second feature allowed if drop in log‑odds ≤ 75th percentile (~0.49).

Outcome:

  • Switching occurs ~20% of days.
  • More diversity in displayed features (people average ~7–8 categories instead of 6).
  • Much fewer extreme long runs.

Still:

  • Some participants remain with long runs (e.g., 40 days) due to genuinely dominant features.

5. Further Ideas Discussed

5.1 Expanding from 13 Categories to All 28 Features

The team strongly favored expanding:

Benefits:

  • More variation → less repetition
  • Feels more “personalized”
  • Messages can be more specific (e.g., “past 24 hours vs. past week”)
  • Better reflection of the data structure
  • More clinical nuance

Risks:

  • Greater precision may reveal model errors
  • But participants can be onboarded to expect occasional inaccuracies

Team consensus:
Move toward 20–28 categories, while still grouping features for intervention mapping.

5.2 Alternative Switching Rules

Ideas captured:

  • Switch based on trend (e.g., if primary feature is improving).
  • Larger runs → increased willingness to switch to another feature.
  • Consider not only drop in predictive value but relative position in daily distribution (quantiles).
  • Use trajectory: “Your top feature is still X, but it’s improving.”

5.3 Feedback Content Enhancements

Proposals:

a. Include trend direction

“Your key risk factor today is craving — but it’s improving compared to yesterday.”

b. Provide both granular and coarse risk categories

Example:

  • “Your recovery score is 87 (High).”

This mirrors credit‑score style messaging.

c. Clarify the role of model error

Explain to participants that:

  • GPS‑based features may be imperfect
  • EMA features are extremely accurate (self‑report)

5.4 Key EDA Needed Before Final Decisions

The team planned exploratory data analysis to examine:

  • Whether long runs correlate with being in very low‑risk or very high‑risk states
  • Differences between participants enrolled for 1 month vs. 3+ months
  • Whether some features dominate simply because of participant profiles (e.g., always angry, or always going to risky places)

6. Final Decisions & Action Items

Modeling Decisions

  • Continue refining v28 with linear PANAS + mins.
  • Improve GPS imputation (NA instead of 0).
  • Expand feature categories possibly to 20–28.

Personalized Messaging Logic

  • Continue evaluating switching rules.
  • Consider quantile‑based or trend‑based switching.
  • Conduct EDA to guide thresholds.

2.4.2 Feedback Design

  • Consider combining numeric and categorical risk displays.
  • Possibly show feature trajectories.
  • Ensure onboarding messaging addresses model imperfection.

NOTE: Interim meetings in this time gap were not about optimize

2.4.3 2025_1211

Discussion centered on CM’s Hilldale project design, LLM-based message generation for Optimize, and strategies to address recruitment fraud in Qualtrics panel studies. Key topics included:

  • Experimental design for tone/style personalization.
  • Baseline preference survey development.
  • Prompt engineering for LLM message generation.
  • Fraud prevention and alternative recruitment strategies.

1. CM’s Hilldale Project & Optimize Context

Background

  • Optimize study uses LLMs to generate daily supportive messages for participants.
  • Messages include:
    • Lapse probability and trend (e.g., high increasing, high decreasing, low increasing).
    • Important risk feature (currently hardcoded as craving).
    • Intervention suggestion.
  • All content wrapped in a supportive tone and style generated by LLM.

Experimental Manipulation

  • Two factors:
    • Style: Formal vs. informal.
    • Tone: Five categories (e.g., legitimizing, norms-based, value affirmation, acknowledging, self-efficacy).
  • Participants:
    • Complete baseline tone/style preference survey.
    • Receive 30 messages (crossing 5 tones × 2 styles × 3 risk scenarios).
      • Down from 48 messages to reduce burden (~15 minutes).
  • Ratings collected for each message:
    • Helpfulness, usefulness, preference, and overall liking.

2. Baseline Preference Survey

Evolution

  • Initial version was complex and abstract (descriptions of tone categories).
  • Simplified approach:
    • Concrete framing:
      “When I need help, I want to be talked to by someone who…”
      Each tone completes the sentence.
  • Goal: Improve clarity and reduce cognitive load.

Open Questions

  • Should survey be administered before or after message exposure?
    • Concern: Participants may not know preferences until they see examples.
    • Idea:
      • Between-subject design: Half complete survey at baseline, half after rating messages.
      • Could reveal whether exposure improves predictive accuracy.

3. Analysis Plan

Three predictive models for message ratings:

  1. Tone/style only:
    • Tests if some tones/styles are globally preferred.
  2. Tone/style + demographics:
    • Examines demographic differences (e.g., age, gender).
    • Implications for equity and personalization.
  3. Tone/style + demographics + baseline preferences:
    • Tests if individual preferences improve prediction beyond demographics.

Interpretation Challenges

  • If preferences do not predict:
    • Could mean no within-group heterogeneity OR survey failed to capture true preferences.
  • Post-hoc option:
    • Analyze within-person variance:
      • Compare variability in ratings within tone/style vs. across tones/styles.
      • If within-tone variance is low and across-tone variance is high → evidence of individual preferences.

4. LLM Prompt Engineering

Structure

  • System prompt:
    • Role: Automated recovery support chatbot.
    • Purpose: Provide daily message with lapse risk info and supportive tone/style.
  • User prompt:
    • Includes:
      • Lapse risk and trend.
      • Important feature (currently craving).
      • Instruction to write in specified tone/style.

Issues & Refinements

  • Overlap between legitimizing and norms-based tones:
    • Both reference “expected part of the process,” leading LLM to include phrases like “many people experience this.”
  • Solutions:
    • Revise tone descriptions:
      • Legitimizing → emphasize validity of feelings without referencing others.
      • Norms → emphasize commonality among peers.
    • Add explicit instruction:
      • “Do not reference other people’s experiences” for legitimizing tone.
  • Consider banned word lists for each tone to prevent cross-contamination.
  • Style currently held constant (formal) for testing; temperature set to 0.7 for variability.

5. Recruitment Fraud Concerns

Problem

  • High fraud rates in online panels (up to 90%+).
  • Risks:
    • Participants misrepresenting eligibility (e.g., claiming alcohol problems).
    • Speeding through surveys or random responding.

Proposed Solutions

  • Mask eligibility criteria:
    • Embed alcohol-related items within broader health/lifestyle screener.
    • Example:
      • Include questions on exercise, diet, mood, and one alcohol question.
      • Use branching logic: Only those meeting alcohol criteria proceed to AUDIT.
  • Two-stage screening:
    • Initial screener with mixed health items.
    • Delay before second stage (if platform allows).
  • Attention checks:
    • Insert “catch” items in message rating task (e.g., obviously inappropriate message).
    • Monitor response time; flag surveys completed too quickly.
  • Explore platform capabilities:
    • Ask Qualtrics about fraud detection tools and multi-step screening feasibility.
  • Alternative recruitment:
    • Local advertising or SONA pool:
      • Pros: Easier verification via phone/Zoom.
      • Cons: Homogeneous sample (college students), potential speeding for credit.
  • Verification ideas:
    • Require short voice recording or Zoom call.
    • Open-ended prompt (e.g., “Talk for one minute about your recovery experience”).
    • IRB considerations for identifiable data.

6. Key Decisions & Next Steps

  • Survey wording:
    • Update legitimizing tone description (use “valid” or “okay” instead of “expected part of process”).
  • Prompt refinement:
    • Add explicit “no reference to others” rule for legitimizing tone.
    • Test banned word approach.
  • Recruitment strategy:
    • Susan to contact Qualtrics about fraud prevention and screening options.
    • Draft screener with mixed health items and branching logic.
  • Attention checks:
    • Design in-context checks for message rating task.
  • Future considerations:
    • Decide whether to administer preference survey post-exposure for subset of participants.
    • Evaluate feasibility of Zoom verification for high-risk cases.

Key Themes

  • Personalization of tone/style may improve engagement, but requires robust measurement and fraud-resistant sampling.
  • LLM prompt engineering must balance clarity and distinctiveness of tone categories.
  • Recruitment fraud is a major threat; multi-layered prevention and detection strategies are essential.

2.4.4 2025_1201

(CAG Meeting)

Emily Dworkin, a licensed clinical psychologist, presented on the growing problem of fraud in online research recruitment, its impact on study validity, and strategies for prevention and detection. The discussion included practical solutions, ethical considerations, and implications for IRB and NIH review.

1. Scope of the Problem

  • Prevalence:
    • Fraud rates in online recruitment are alarmingly high:
      • 82% (Bell & Gift, 2022)
      • 92% (Johnson et al., 2023)
      • 99% (Shaw et al., 2025)
    • Trend worsened during COVID and continues to escalate.
  • Affected platforms:
    • Mechanical Turk, social media ads, message boards, listservs.
  • Consequences:
    • One investigator had to shut down a study and personally repay NIH after discovering 10,000 fraudulent participants.
  • Motives:
    • Financial gain: Particularly impactful in regions where small incentives are significant.
    • Mischief or sabotage: Politicized topics attract bad actors aiming to distort findings.

2. Types of Fraud

  • Bots:
    • Least concerning; relatively easy to detect and block.
  • Human fraud:
    • Falsifying eligibility.
    • Participating under multiple identities.
  • Hybrid schemes:
    • Bots and humans working together (e.g., survey farms).

3. Harms

  • Wasted staff time and resources.
  • Incentives diverted from legitimate participants.
  • Invalid data compromising study outcomes.
  • Risk of funder repayment and reputational damage.

4. Prevention Principles

Verification of Eligibility

  • Identify which criteria can be verified and implement checks:
    • Age:
      • Ask at multiple time points in different formats; compare for consistency.
    • Location:
      • Use metadata and/or Zoom verification.
      • IP address alone is unreliable.
      • On Zoom:
        • Ask participant to show phone screen (time zone/weather app).
        • Request live location or quiz on local details (area code, nearest highway, county name).
    • Device/browser settings:
      • Collect metadata at every survey step.
    • Demographics:
      • Ask immutable characteristics (e.g., race) multiple times.
    • Insider knowledge:
      • Use population-specific questions (e.g., military unit for veterans).
      • Caveat: Can be spoofed by AI tools like ChatGPT.

Unverifiable Criteria

  • Examples: sexual orientation, gender identity, symptom severity.
  • Mask these criteria:
    • Do not disclose in ads or consent forms.
    • Example: Emily’s study targeted sexual minority women but framed recruitment as a stigma study.
    • Use secret identifiers (e.g., “WLW”) known only to the target group.

5. Process Design

  • Apply same screening process for all participants:
    • No skip logic.
    • Do not inform participants of ineligibility.
  • Communicate that all screeners will be contacted if selected.
  • Introduce delays between screening and follow-up:
    • Emily used an automated 5-minute delay before next steps.
  • Multi-tiered approach:
    • Preliminary screener → second screener → waitlist message.
    • Metadata checks at every stage (device, browser, IP, geolocation).
    • Staff review before baseline survey.
    • Zoom verification for suspicious cases:
      • Show ID (front and back).
      • Share phone screen or answer location quiz.
      • Screen sharing for live location (mobile feasibility unclear).
      • High school mascot question (hard to fake quickly).

6. Detection Outcomes

  • Initial pool: 25,000 prescreened.
  • Second screen: 17,000.
  • Invited: 3,800.
  • Baseline completed: 2,612.
    • 95% legitimate.
    • 5% mistakenly ineligible.
    • 3% excluded for likely fraud.
  • Zoom calls required for <10% of participants.
  • Impact:
    • Fraudulent responses would have significantly altered primary outcomes (e.g., trauma history, PTSD severity, AUDIT scores).

7. Ethical & Practical Considerations

  • Transparency:
    • Explain verification steps as protecting participant contributions and study integrity.
  • IRB challenges:
    • Some IRBs resist collecting extra data; Emily’s IRB was supportive.
    • Draft language available for consent and IRB submissions.
  • Confidentiality:
    • Verification lookups (e.g., phone number checks) are low risk compared to storing sensitive data.
  • Recruitment budgets:
    • Must account for fraud prevention costs.

8. Platform Limitations

  • Qualtrics and REDCap:
    • Basic checks (cookies, IP, VPN detection) insufficient.
    • Do not cross-check PII against external databases.
  • Emily developing an API tool to integrate robust fraud detection with these platforms.

9. NIH & IRB Implications

  • NIH study sections increasingly scrutinize fraud prevention rigor.
  • Reviewers expect evidence of robust detection in published research.
  • Responsibility extends beyond participant safety to ensuring valid interpretation of their contributions.

Key Takeaways

  • Fraud in online recruitment is pervasive and escalating.
  • Prevention alone is insufficient; multi-layered detection is essential.
  • Masking eligibility criteria and verifying immutable data points are critical.
  • Robust fraud prevention should be standard practice and explicitly documented in protocols and publications.

2.4.5 2025_1120

Discussion centered on emotion measurement for Optimize EMA, feature engineering for predictive models, and GPS context questions. Key topics included:

  • Moving away from discrete emotion measures toward valence/arousal dimensions.
  • Improving participant understanding of affect ratings.
  • Timing and framing of EMA questions.
  • Refining GPS context questions for Optimize based on lessons from Risk1 and Risk2.

1. Emotion Measurement: Valence & Arousal vs. Discrete Emotions

Challenges with Discrete Emotions

  • Discrete emotion approach (e.g., PANAS):
    • Requires many adjectives (20–34 items) for adequate coverage.
    • Short versions lose precision; descriptors often fail to capture participant experience.
    • Example: “Afraid” vs. “Angry” both high-arousal negative affect, but participants may reject one even if they fit the quadrant.
  • High burden for EMA (multiple times daily) makes this impractical.

Decision

  • Stay with valence and arousal dimensions (two-item approach):
    • Lower dimensional representation predicts much of variance in discrete emotions.
    • Easier to implement and less burdensome.
  • Goal: Improve clarity and participant comprehension.

Improving Measurement

  • Psychoeducation at intake:
    • Show participants a visual affect grid or circumplex model.
    • Explain dimensions (pleasant ↔︎ unpleasant; high ↔︎ low activation).
    • Provide examples of adjectives in each quadrant.
  • Anchors:
    • Need better labels for high/low ends of scales.
    • Susan tasked to:
      • Gather examples of affect grids and circumplex models.
      • Collect commonly used dimension labels and anchors.
      • Find visuals for participant training and in-app reference.
  • Implementation ideas:
    • Display small reference image during EMA.
    • Explore mood meter concepts (e.g., “How We Feel” app).
    • Consider sequential questions (e.g., Finch app approach):
      • First: Pleasant/unpleasant.
      • Then: High/low energy.
      • Possibly provide adjective examples based on initial response.

Feature Engineering Considerations

  • Include:
    • Valence term.
    • Arousal term.
    • Interaction term (valence × arousal).
    • Quadratic terms for non-linear effects (e.g., extremes of valence/arousal predict higher risk).
  • Rationale:
    • Risk may depend on quadrant (e.g., high-arousal negative affect).
    • Current GLMNET models lack interactions/quadratics; XGBoost allows them but not interpretable.
  • JC plans:
    • Refit models on Risk1 data to test interactions and quadratic effects.
    • If significant, incorporate these features into Optimize models.

2. Timing of EMA Affect Questions

  • Current Optimize design: Evening EMA (end of day).
  • Debate:
    • Right now vs. retrospective (average of day):
      • “Right now” aligns with prediction model but may reflect bedtime calm rather than daytime risk.
      • Retrospective captures overall emotional tone but risks ambiguity (multiple emotions during day).
  • Prospective (tomorrow’s expected mood) considered but likely less accurate for affect.
  • Decision pending:
    • Optimize may keep “right now” for prediction compatibility.
    • Consider adding retrospective question for future modeling.

3. GPS Context Questions

Risk1 vs. Risk2 Comparison

  • Risk1:
    • Focused on risk only (high/medium/low).
    • Did not capture protective aspects or nuanced context.
  • Risk2:
    • Broke context into multiple questions:
      • Type of place (home, shelter, restaurant, etc.).
      • Activities at location (select all that apply).
      • Emotional valence (pleasant/unpleasant).
      • Recovery impact (helpful/harmful).
      • Alcohol availability.
  • Issues in Risk2:
    • Participants often selected many activities, sometimes contradictory (e.g., “drink alcohol” + “get health care”).
    • High frequency of “I don’t recognize this location” responses (≈36%).
      • Possible reasons:
        • Stoplights or transient locations detected.
        • Participants avoiding sensitive disclosures.
        • Burden reduction (skipping detailed questions).

Proposed Changes for Optimize

  • Keep Risk2 structure but:
    • Correct skip logic (ensure all relevant questions asked for home and temporary residences).
    • Add:
      • “Alcohol available here?”
      • “Have you ever drunk alcohol here?”
    • Limit activity selection:
      • Consider top 3 activities instead of “select all that apply.”
  • Passive data integration:
    • Use geocoding and public data to infer type (e.g., restaurant, business).
    • Still ask participants what they do there for accuracy (same location can have different meanings for different people).
  • EDA planned:
    • Analyze Risk2 data for:
      • Frequency of multi-select activities.
      • Patterns in “I don’t recognize” responses.
      • Free-text comments for context.

Stoplight Issue

  • Risk2 used 5-minute stationary threshold for “important place.”
  • Still saw many transient locations flagged.
  • Possible solutions:
    • Increase duration threshold.
    • Validate against geocoded place types.
    • Consider requiring multiple visits for classification.

4. Recruitment Scope

  • JC leaning toward national sample for Optimize:
    • Faster recruitment.
    • Greater diversity for model improvement.
    • Lower burden than coordinating with Wisconsin treatment providers.
  • Integrity measures:
    • Implement automated checks to prevent fraudulent participation.
    • Lessons learned from Risk2 inform screening protocols.

Next Steps

  • Susan:
    • Draft updated GPS context question set (based on Risk2, corrected skips, added alcohol questions).
    • Collect affect grid examples and dimension anchors.
  • Kendra:
    • Conduct EDA on Risk2 GPS data:
      • Multi-select patterns.
      • “I don’t recognize” frequency and comments.
  • JC:
    • Test interaction and quadratic effects for valence/arousal on Risk1 data.
    • Explore timing implications for EMA affect questions.
  • Team:
    • Decide on activity selection threshold (e.g., top 3).
    • Finalize recruitment strategy (national vs. regional).

Key Takeaways

  • Emotion measurement will remain valence/arousal-based, with improved participant guidance and feature engineering.
  • Optimize EMA design will:
    • Retain predictive compatibility.
    • Add retrospective or prospective items for future modeling.
  • GPS context questions will:
    • Capture both risk and protective factors.
    • Reduce ambiguity and participant burden.
  • Recruitment likely national for scalability and diversity.

2.4.6 2025_1113

Discussion focused on EMA design for Optimize, feature retention for predictive modeling, and improvements to emotion measurement frameworks. Key topics included:

  • Alignment of EMA items with Risk1 and Risk2 models.
  • Decisions on adding or dropping questions.
  • Conceptualization of emotional affect dimensions for better feature engineering.

1. Infrastructure Update

  • Docker deprecated:
    • All team members must use VM environment.
    • Geocoding database will be downloaded locally → no need for Docker.

2. EMA Design & Model Compatibility

Framing

  • Optimize must support predictions using Risk1 model:
    • Requires inclusion of all features used in Risk1 model.
  • Question: Can we drop features not retained in final model?
    • Answer: Yes, because GLMNET models only need retained features for predictions.
  • Why add new features?
    • To train new models on Optimize data (larger sample, longer duration).
    • Potential to improve predictive accuracy and intervention logic.
    • Could integrate Risk1 data with Optimize for future modeling.

JC Decisions

  • Additions:
    • Bring back future risky and future stressful events.
    • Add future pleasant events.
  • Retention:
    • Nothing dropped from Risk1 for prediction compatibility.
  • New items:
    • Keep probability of future drinking (from Risk1).
    • Add confidence and motivation questions (similar to Risk2).

Risk1 vs Risk2 EMA Comparison

  • Risk1:
    • Probability of future drinking.
    • Motivation-like question (only collected post-lapse).
    • Emotion measured as valence/arousal; only arousal retained.
      • High arousal = protective (calm interpreted as boredom).
  • Risk2:
    • Separate confidence and abstinence motivation items.
    • Emotion measured via five discrete emotions:
      • Depressed (unpleasant, low arousal).
      • Angry (unpleasant, high arousal).
      • Anxious (unpleasant, high arousal).
      • Relaxed (pleasant, low arousal).
      • Happy (pleasant, high arousal).
    • Stressful & pleasant events measured for past only (no future).

3. Emotion Measurement Framework

Current Issues

  • Valence/arousal framework theoretically strong but confusing for participants.
  • Most SUD risk linked to negative affect; need better quadrant representation.

Proposed Quadrant Model

  • Axes:
    • Valence: Pleasant ↔︎ Unpleasant.
    • Arousal: Low activation ↔︎ High activation.
  • Quadrants:
    • Activated Positive (e.g., elated).
    • Unactivated Positive (e.g., calm, content).
    • Activated Negative (e.g., angry, anxious).
    • Unactivated Negative (e.g., sad, depressed).
  • Neutral zone:
    • Midpoint between axes (score near zero).

Implementation Ideas

  • Use two core questions (pleasantness & arousal) to infer quadrant.
  • Feature engineering:
    • Count EMA responses per quadrant for last day, 3 days, week.
  • Psychoeducation:
    • Show quadrant map during consent to aid comprehension.
  • Alternative:
    • Heat map in Qualtrics (similar to “How We Feel” app).
  • Consider dominance dimension:
    • E.g., Enraged (high dominance) vs Anxiety (low dominance).

Intervention Implications

  • Quadrant-level risk may guide intervention type.
  • KP: For clinical application, focus on specific emotions, not just quadrant.
  • JC: Explore whether quadrant-based features improve prediction vs raw valence/arousal.

4. Next Steps

  • Feature engineer emotion data using valence/arousal and quadrant approach.
  • Validate whether combined features outperform original measures.
  • Add confidence and motivation items to EMA battery.
  • Continue refining emotion framework for Optimize.

Key Takeaways

  • Optimize EMA will:
    • Retain Risk1 features for prediction compatibility.
    • Add new items for confidence, motivation, and future events.
  • Emotion measurement:
    • Move toward quadrant-based representation for better interpretability and predictive power.
  • Intervention design:
    • Consider mapping risk categories to actionable strategies.

2.4.7 2025_1106

Finalizing feature importance and category-based intervention logic for the Optimize project, addressing:

  • Model calibration issues.
  • Selection of predictive features.
  • Grouping features into actionable categories.
  • Strategies for recommending recovery modules based on daily risk predictions.

1. Feature Importance & Model Calibration

SHAP Limitations

  • SHAP values work well for feature importance only if models are well-calibrated.
  • Problem:
    • Optimize models (raw XGBoost) were poorly calibrated.
    • Result: Narrow probability range (e.g., 0.48–0.52), making SHAP values small and similar across features.
  • Calibration challenge:
    • SHAP for XGBoost requires native model; recalibrated models need custom SHAP computation (impractical for large datasets).
  • Implication:
    • Current feature importance reports are based on uncalibrated models, which may misrepresent true predictive influence.

Calibration Observations

  • EMA-only models: Reasonably calibrated.
  • EMA + GPS models: Poor calibration; even after adjustment, improvement was limited.
  • Calibration patterns:
    • Should be monotonic (higher predicted risk → higher actual lapse probability).
    • Observed quadratic patterns in some cases.
  • Outstanding question: Why certain data characteristics lead to poor calibration.

2. Model Selection Decisions

  • Pivoted from XGBoost to GLMNET (regularized logistic regression):
    • GLMNET outputs probabilities and is generally well-calibrated.
    • Performance trade-off minimal (AUROC difference ≈ 0.01).
  • Best-performing GLMNET models:
    • Lasso (alpha = 1) retained only 3–4 features → too little diversity.
  • Final choice:
    • Models with AUROC ≈ 0.91, less regularized (lower alpha, lower lambda).
    • Retained more features for interpretability and diversity.

3. Coefficient Interpretation

  • Positive coefficient = higher lapse risk as feature increases.
  • Retained raw EMA/GPS features (excluded difference-from-baseline scores for clarity).
  • Key retained features:
    • EMA1: Confidence (commitment to abstinence).
    • EMA10: Future drinking likelihood.
    • Craving.
    • EMA3: Risky situations.
    • EMA4: Past stressors (negative in model; positive in bivariate correlation).
    • EMA7: Arousal (negative; boredom linked to higher risk).
    • EMA8: Future risky situations.
    • GPS: Time at avoid locations, unpleasant places (protective), work (risk-enhancing).
    • Day of week: Monday protective, Saturday high risk.

Direct vs. Total Effects

  • Issue: Coefficients reflect direct effects controlling for other features, which complicates interpretation.
  • Solution:
    • Compute bivariate correlations for total effects using corrr package.
    • Observations:
      • EMA4 positively correlated with lapse risk overall.
      • Time at unpleasant places remains protective.
      • EMA7 negative correlation (boredom → higher risk).

4. Feature Grouping into Categories

Grouped correlated features for category-level predictions and interventions:

Categories

  • Past Use:
    • Lapse history across windows.
  • Commitment to Abstinence:
    • EMA1 (confidence).
  • Craving:
    • Single EMA feature.
  • Risky Situations/Locations:
    • EMA3, EMA8, GPS avoid/risk-medium.
    • Challenge: EMA and GPS risk features weakly correlated.
    • Discussion:
      • Should GPS avoid time be separate? Indicates recovery plan failure.
      • Medium-risk GPS locations ambiguous; correlation low.
  • Stress:
    • EMA4 (past stressors) + GPS unpleasant places.
    • Mixed directionality (EMA positive, GPS negative).
  • Activated:
    • EMA7 (arousal).
  • Future Drinking Probability:
    • EMA10.
  • AA Attendance:
    • Standalone predictor.
  • Day of Week:
    • Monday, Sunday protective; Saturday high risk.
  • Lifestyle:
    • High variability in locations, restaurants, community spaces, work.
    • Currently weak category; needs refinement.

Intervention Considerations

  • Categories should align with actionable recommendations:
    • Risky locations → planning strategies.
    • Avoided places → address recovery plan adherence.
    • Stress → coping skills.
  • GPS complexity:
    • Some “unpleasant” places (e.g., therapy, gym) are protective despite negative valence.
  • Future work:
    • Allow features to belong to multiple categories if intervention logic overlaps.

5. Daily Recommendation Logic

  • For each participant/day:
    • Primary method: Select category with highest predicted lapse risk.
    • Secondary method:
      • Consider second-ranked category if:
        • Risk difference ≤ 25th percentile of first.
        • Avoid repeating same category > 3 consecutive days.
      • Skip second category if drop > 50th percentile.
  • JC leaning toward option 2 for diversity in recommendations.

6. Key Insights

  • Past use rarely surfaces for non-lapsers (good; limited intervention potential).
  • AA appears mainly for low-risk individuals (risk driven by non-attendance).
  • Lifestyle category underdeveloped; needs conceptual clarity.
  • GPS-based recommendations should emphasize reflection on time spent in locations.

Next Steps

  • Refine category definitions and intervention mapping.
  • Reassess GPS feature correlations across full feature set (≈300 features).
  • Explore multi-category membership for features.
  • Prepare updated logic for daily module recommendations.
  • Revisit in two weeks with proposed refinements.

Discussion Highlights

  • Calibration issues remain an open research question.
  • Importance of distinguishing direct vs. total effects for interpretability.
  • Intervention design must consider behavioral context, not just statistical correlation.

2.4.8 2025_1009

Review and refine recruitment materials, website content, and study description for the Optimize project, incorporating feedback from CARDS advisory boards. Discussion included:

  • Privacy and data use messaging.
  • Visual design and branding consistency.
  • Language clarity and tone for recruitment and consent.
  • Strategies for participant engagement and transparency.

Key Discussion Points

1. Privacy & Data Use

  • CARDS feedback: Participants were interested in knowing their data could help improve future models.
  • Action:
    • Include a statement emphasizing that participant feedback and data will improve future versions of the system, not real-time personalization during the study.
    • Remove or rephrase language suggesting participants can “flag bad predictions” (misleading; cannot distinguish between disagreement and true inaccuracy).
  • Simplification:
    • Current privacy section is text-heavy; consider breaking into smaller pages or using bold headings and bullet points.
    • Provide optional links for detailed documents (e.g., Certificate of Confidentiality) for those who want more information.

2. Study Description & Messaging

  • Current framing risks implying the system is a replacement for treatment.
  • Revised focus:
    • Emphasize continuing care and personalized support, not treatment substitution.
    • Highlight dynamic nature: “Tools designed to fit into daily life and adapt to changing needs.”
  • Personalization language:
    • Downplay predictions in recruitment materials to avoid confusion for participants in non-personalized conditions.
    • Focus on: “Your survey responses help tailor supportive messages” rather than “personalized predictions.”
  • Future orientation:
    • Clarify that personalization and predictive modeling are long-term goals, not guaranteed improvements during the study.

3. Branding & Visual Design

  • Consistency issue:
    • Website uses UW red theme; flyers use older Addiction Research Center branding.
  • Consensus:
    • Use UW-Madison branding prominently for credibility.
    • Include both UW logo and ARC logo on recruitment materials.
  • Logo terminology:
    • Debate over “Addiction Research Center” vs. softer language (e.g., “Recovery Research”).
    • Decision: Keep ARC name for transparency; avoid stigmatizing language in participant-facing text.
  • Visual updates:
    • Current designs look dated; need modern, clean graphics.
    • Suggestion: Use infographics and icon-based layouts for privacy and FAQ sections.
    • Explore NIH “All of Us” study as a model (image-based consent and recruitment materials).
  • Action:
    • Consider hiring a graphic designer or using stock images (neutral, relatable, diverse).
    • Avoid images showing alcohol or cigarettes (trigger risk).
    • Use photos of people on phones, neutral or positive expressions.

4. Recruitment Materials

  • Formats:
    • Posters for clinics/community centers.
    • Two-sided brochures for detailed info.
    • Business cards for word-of-mouth referrals.
    • Digital ads for Craigslist and Reddit (previously effective).
  • Content priorities:
    • Must include:
      • UW affiliation.
      • Statement that it is research.
      • Compensation details (without overemphasis per IRB rules).
    • Avoid language implying “cutting-edge treatment”; instead use “help advance research.”
  • Enhancements:
    • Add infographic of study timeline (screening → visits → surveys).
    • Include testimonial quotes from prior studies (positive feedback from burden survey).
    • Tagline encouraging participants to share with friends/family.
  • Word-of-mouth strategy:
    • Provide participants with cards or links for easy sharing.
    • IRB likely prohibits paid referral bonuses; unpaid referrals encouraged.

5. Website Structure

  • Landing page should:
    • Summarize study purpose in plain language.
    • Offer quick access to screening survey.
    • Include FAQs and optional deep-dive links for privacy/legal details.
  • Consider adding:
    • Video introduction (e.g., John explaining study goals).
    • Section for research transparency:
      • Publications with plain-language summaries.
      • How participant feedback informs improvements.

6. Language Consistency

  • Use phrases like:
    • “Recently quit drinking” or “in recovery from past problems with alcohol use.”
  • Avoid overly clinical terms (e.g., “alcohol use disorder”) and stigmatizing language.
  • Maintain consistency across all materials (website, flyers, consent).

7. Participant Engagement Ideas

  • Testimonials:
    • Use published quotes from burden survey and prior studies.
  • Transparency:
    • Optional section showing how feedback shapes future research.
  • Video content:
    • Short interview with John (2–3 questions) for website.
    • Possibly link his TED Talk for context on digital therapy innovation.

8. CARDS Review & Timeline

  • Two virtual meetings scheduled:
    • Nov 24 (Lucier group).
    • Dec 16 (Goodman group).
  • Materials due three weeks prior to first meeting.
  • Deliverables for CARDS:
    • Recruitment materials (poster, brochure, business card, digital ads).
    • Website mock-up with study description, FAQs, and team bios.
    • Screening survey and FAQ for staff.
    • Draft onboarding script.

Action Items

  • Susan:
    • Revise recruitment materials with updated branding and simplified language.
    • Prepare website shell with tabs for study info, FAQs, and team bios.
    • Collect testimonial quotes and draft infographic.
  • Team:
    • Submit bios and headshots for website.
    • Suggest stock images or design ideas.
  • John:
    • Approve final language for study description and consent consistency.
    • Confirm personalization messaging approach.
  • Next Meeting:
    • Focus on visual design and recruitment material revisions.
    • Review updated website structure and infographic drafts.

Key Themes

  • Clarity and trust: Simplify language, emphasize UW affiliation, and highlight research purpose.
  • Modern design: Use visuals and infographics to reduce cognitive load.
  • Participant-centered approach: Offer transparency, testimonials, and optional deep-dive info.
  • Consistency: Align terminology across all materials and platforms.

2.4.9 2025_1002

Planning for CARDS advisory board review, recruitment strategies, and participant-facing materials for the Optimize study. Discussion centered on:

  • Website and recruitment materials.
  • Screening process and SOPs.
  • Participant onboarding and consent scripts.
  • Timeline for CARDS meetings and deliverables.

Key Discussion Points

1. Public Persona & Professional Presence

  • Emphasis on maintaining a professional social media presence for future careers.
  • Increasing importance of public visibility for researchers and clinicians.
  • No clear separation between private and professional life in the current digital environment.

2. CARDS Feedback & Website Design

  • CARDS strongly recommended adding a “Meet the Team” section to build trust and transparency.
  • Historically, research teams avoided personal profiles, but participants want to know who is behind the study.
  • Proposed structure:
    • Landing page with study overview, FAQs, and contact info.
    • Tabs for:
      • Study details.
      • Frequently Asked Questions (FAQ).
      • Team bios and photos.
  • Team bios should convey:
    • Expertise relevant to the study.
    • Why the researcher cares about the project.
    • Trust-building elements.
  • Decision pending on how many team members to include (avoid overwhelming participants).
  • Consider highlighting roles tied to data-sensitive tasks (e.g., GPS handling, LLM message generation).

3. Recruitment Materials

  • Types of materials:
    • Local:
      • Posters for clinics/community centers.
      • Brochures (two-sided, detailed).
      • Business cards for word-of-mouth referrals.
    • Digital:
      • Craigslist and Reddit posts (previously effective).
      • Explore TikTok or other social media (likely low priority given target demographic).
    • Website:
      • Central hub for study info, FAQs, and screening link.
      • QR codes and short URLs for easy access.
  • Word-of-mouth strategy:
    • Participants encouraged to share study info.
    • IRB concerns about paid referral bonuses; likely not permitted.
    • Alternative: provide extra business cards or digital links for sharing.
  • Content considerations:
    • Clear, concise language.
    • Include QR codes and short URLs.
    • Infographic showing study timeline and tasks (screening → visits → surveys).

4. Screening Process

  • Primary mechanism: Online Qualtrics screening survey.
  • Phone option:
    • Available for participants uncomfortable with online forms.
    • Staff will answer questions but ultimately direct participants to complete the survey themselves.
  • Rationale:
    • Participants must be comfortable with digital tools for EMA compliance.
  • FAQ for staff:
    • Comprehensive document with answers to common participant questions.
    • Updated continuously based on new queries.
    • Will be shared with CARDS for feedback.

5. Onboarding & Consent

  • Onboarding script:
    • Critical for participant trust and clarity.
    • Will include:
      • Study purpose.
      • Data privacy and security.
      • Explanation of tasks and compensation.
  • Consent form:
    • Already IRB-approved for STAR system.
    • Updates needed:
      • Replace references to standalone app with “STAR system.”
      • Clarify use of FollowMee for GPS and text messaging for surveys.
  • Financial disclosure:
    • Simplify language regarding Dave Gustafson’s historical association with CHESS.
    • Avoid unnecessary IRB complications.

6. Study Visits

  • Three main visits:
    • Baseline: Consent, Timeline Followback, initial survey.
    • Two-month: Timeline Followback + check-in.
    • Four-month: Timeline Followback, exit survey, debrief, and optional qualitative interview.
  • Surveys delivered via Qualtrics links, not administered live during visits.
  • Exit interviews:
    • Will gather feedback on system usability, costs/benefits, and participant experience.
    • May require separate CARDS review.

7. CARDS Meeting Timeline

  • Proposed dates:
    • Lucier Group: November 24 (preferred).
    • Goodman Group: December 16 (backup).
  • Materials due three weeks prior to meeting.
  • Target: Finalize recruitment materials and website by November 1.

8. Action Items

  • Susan:
    • Draft recruitment materials (poster, brochure, business card, digital ads).
    • Build website shell with tabs for study info, FAQs, and team bios.
    • Prepare infographic for study timeline.
  • Team Members:
    • Submit short bios and headshots:
      • Highlight expertise and motivation.
      • Keep tone trust-building and participant-friendly.
  • John:
    • Review screening survey and finalize onboarding script.
    • Coordinate with CARDS on review scope and expectations.
  • Next Meeting:
    • Focus on recruitment materials (local, digital, website).
    • Review draft bios and website structure.

Strategic Considerations

  • Recruitment approach must balance digital reach with local engagement.
  • Transparency and trust are key:
    • Bios and team profiles should humanize the study.
    • Clear communication about data handling and privacy.
  • IRB compliance:
    • Avoid referral payments; consider alternative engagement strategies.
  • CARDS input will guide:
    • Material design.
    • Language clarity.
    • Cultural sensitivity.

Summary of Deliverables for CARDS

  • Recruitment materials:
    • Poster, brochure, business card, digital ad mockups.
  • Website landing page with:
    • Study overview.
    • FAQs.
    • Team bios.
  • Screening survey link and FAQ for staff.
  • Draft onboarding script and consent updates.

2.4.10 2025_0925

Finalize Optimize study materials and address key issues related to:

  • Baseline measures and survey synchronization.
  • Screening criteria and participant eligibility.
  • IRB-approved consent language and disclosures.
  • Preparation for CARDS review and NDA submission.

1. Baseline Measures & Survey Synchronization

  • Issue: Discrepancy between demographic survey, QMD files (ground truth), Qualtrics survey, and Word document.
  • Resolution:
    • QMD files will serve as authoritative source for all measures.
    • Word document will retain comments for review but will not be primary reference.
    • Susan to update Word doc to match QMD and transfer comments for record.
  • Action:
    • John to finalize decisions on measures using QMD + comments.
    • Ensure Baseline battery includes:
      • Individual difference measures.
      • First administration of repeated outcome measures (except system-specific ones).
  • Validated measures (e.g., PHQ-9, WHO ASSIST) will remain unchanged; reading level adjustments apply only to non-validated content.

2. Monthly Measures

  • Discussion:
    • Recovery Capital (MIRC) confirmed as repeated measure and secondary outcome.
    • No monthly predictors currently used in the model; likely removed from predictive algorithm.
    • Monthly measures retained for future modeling and sample characterization.
  • Outcome:
    • Optimize will include repeated measures for outcomes but not predictors.

3. Screening Criteria & Eligibility

  • Clarifications:
    • Smartphone requirement:
      • Remove “newer smartphone” language; any smartphone with data plan acceptable.
      • Data plan assumed standard; handle exceptions case-by-case.
    • Recovery status:
      • Drop requirement for prior formal treatment.
      • Focus on recently quit drinking (within last 4 months) and goal for future alcohol use.
  • Revised Screening Questions:
    • “Have you recently quit drinking alcohol?”
    • “What is the date of your last drink?”
    • “What is your goal regarding future alcohol use?” Options:
      • I don’t want to drink at all.
      • I want to cut down on frequency (days per week).
      • I want to cut down on amount per occasion.
      • I don’t yet have a specific goal.
  • Other Adjustments:
    • Remove references to video check-ins and standalone app.
    • Simplify language for clarity and reduce reading level (target grade 5–6).

4. WHO ASSIST & Substance Use Questions

  • Current Issue:
    • Questions framed as lifetime use; follow-ups may not align with recent behavior.
  • Decision:
    • Initial question: “Which substances have you used for non-medical reasons in the past 3 months?”
    • Follow-up questions:
      • Craving and problems → retain lifetime framing for sensitivity.
  • Rationale:
    • Past 3-month use better characterizes current risk; lifetime questions remain for context.

5. IRB & Consent Language

  • Changes Needed:
    • Update consent and participant-facing materials:
      • Replace “standalone app” references with “STAR system” or “study tools.”
      • Clarify use of FollowMee for GPS and text messaging for surveys.
    • Financial disclosure for Dave Gustafson:
      • Current language unclear; revise for transparency without triggering unnecessary IRB review.
  • Strategy:
    • Avoid changes that could require NIH program officer approval unless essential.
    • Frame updates as clarifications for participant understanding.

6. NDA Submission

  • Deadline: October 1, 2025.
  • Requirements:
    • Generate GUIDs and pedigrees using real birth information for pilot participants.
    • Collect baseline data ASAP to populate NDA.
  • Plan:
    • Susan to resend birth info form and updated baseline survey.
    • Team to complete forms promptly; John to finalize measure decisions by end of week.

7. CARDS Review Preparation

  • Materials for Review:
    • Screening questions.
    • Recruitment ads and participant-facing descriptions.
    • Study introduction script.
    • Consent form (cannot change but provide for context).
  • Goal:
    • Ensure clarity, trust-building, and cultural sensitivity in recruitment and onboarding.

8. Participant Communication & Website

  • Discussion:
    • Create a study landing page for participant info and recruitment.
    • Include:
      • Study overview.
      • Screening link.
      • Contact details.
    • Recovery activity content will be served separately (not public).
  • Security:
    • No password protection needed for general info page; screening link can be public.

9. Reading Level & Accessibility

  • Target: Grade 5–6 for all non-validated content (study descriptions, messages, recovery modules).
  • Tools:
    • Grammarly, Hemingway, and R-based readability analysis.
  • Validated measures: No changes to wording.

Action Items

  • John:
    • Finalize baseline measures and QMD updates by tomorrow.
    • Review comments and resolve decisions.
  • Susan:
    • Sync Word doc with QMD.
    • Send birth info form and updated baseline survey to team.
    • Begin drafting CARDS materials (ads, scripts, study description).
  • Team:
    • Complete birth info and baseline survey promptly.
  • Future Tasks:
    • Prepare recruitment website.
    • Confirm scheduling process (likely phone-based; avoid online calendaring for privacy).

Key Themes

  • Streamlining measure management and ensuring consistency.
  • Simplifying screening and consent language for clarity and trust.
  • Balancing IRB compliance with practical implementation.
  • Preparing for CARDS input and NDA submission under tight timelines.

2.4.11 2025_0904

Topic: GPS Data Processing and Feature Extraction for Optimize Project

Presenter: Chris

Objective

Explore methods to derive meaningful behavioral and environmental features from uncontextualized GPS data to:

  • Reduce reliance on manually provided context.
  • Potentially capture subconscious patterns in participant routines.
  • Ensure privacy while maximizing utility for predictive modeling.

Proposed Workflow

  1. Classify Movement State:
    • Distinguish stationary vs. movement.
    • Focus on stationary periods as “meaningful activity stops.”
  2. Cluster Locations:
    • Group GPS points into clusters representing routine places (e.g., home, work).
    • Challenge: FollowMee pings are sparse; standard clustering fails → use time-duration method (≥30 min threshold).
  3. Enrich with Metadata:
    • Reverse geocoding to retrieve contextual info:
      • Business type.
      • Alcohol availability.
      • Neighborhood characteristics.
      • Proximity to high-risk environments.
    • Potential to integrate Madison zoning codes for additional context.
  4. Derive Features:
    • Time spent at different location types.
    • Stability of housing and work routines.
    • Exposure to high-risk environments (e.g., bars, liquor stores).
    • Granularity guided by privacy vs. utility trade-off.

Technical Infrastructure

  • Tools & Frameworks:
    • PostGIS: Store and represent geolocation data (cannot be handled in CSV).
    • Docker VM: Two containers:
      • PostGIS for analysis (recommended by Rundle; sophisticated).
      • Nominatim (OpenStreetMap) for geocoding (excellent Wisconsin coverage; WI file ~30GB).
    • R: Data manipulation and processing.
    • SQL: For memory-intensive operations.
  • Advantages:
    • Fully reproducible pipeline.
    • Mounted on restricted research drive → no external data transfer.
  • Third-party geocoding:
    • ARC-GIS geocoder deemed inadequate.
    • Nominatim chosen for local coverage and privacy control.

Privacy & Compliance Discussion

  • Data Sensitivity Categories:
    • IID (Individually Identifiable Data).
    • PII (Personally Identifiable Information).
    • PHI (Protected Health Information; HIPAA-regulated subset of PII).
  • Key Points:
    • Our GPS data generally not health data, but linked to AUD research → gray area.
    • Collaborators in Health Care Components (HCC) (e.g., Randy Brown) → may classify data as PHI.
    • Rundle’s stance: A single GPS point = IID health data if tied to AUD research.
    • JC’s position: A lone coordinate without time or subject ID is not identifiable.
  • Risks:
    • Reverse geocoding queries could expose residential locations.
    • Logs of queries by third-party services.
    • Publishing aggregated location info could inadvertently reveal population details.
  • Mitigation:
    • Batch process all locations across participants → millions of queries, anonymized.
    • Keep raw GPS data on restricted drive; only derived feature-level files on shared drive.
    • Use VPN and internal servers to reduce external exposure.

Next Steps

  1. Finalize and document Docker-based geocoding pipeline.
  2. Chris to meet with John to review clustering methods (compare John’s code vs. alternatives).
  3. Ensure restricted drive storage for raw GPS data; preprocess before moving to shared drive.
  4. Susan to provide Chris access to Linux VM for implementation.

Key Takeaways

  • Moving toward feature-level representation of GPS data for privacy and scalability.
  • Combining technical rigor (PostGIS, Docker, reproducibility) with privacy safeguards.
  • Focus on extracting behaviorally meaningful features (routine stability, risk exposure) for predictive modeling.

2.4.12 2025_0821

Main Objective

Design and refine a qualitative interview guide for participants in the Optimize study to gather in-depth feedback on:

  • User experience with the monitoring and support system.
  • Costs, benefits, and barriers of participation.
  • Perceptions of algorithm-driven predictions and LLM-generated messages.
  • Identify gaps in current quantitative measures and inform future improvements.

Key Goals of the Exit Interview

  1. Stakeholder Engagement:
    • Treat participants as stakeholders to understand what they value and what would make the system more useful.
    • Gather insights for future grants and real-world implementation.
  2. Mixed-Methods Approach:
    • Complement quantitative measures (e.g., digital working alliance scales) with qualitative data.
    • Explore discrepancies between what participants say and what current measures capture.
  3. Focus Areas:
    • Algorithm Inputs: EMA responses, GPS data, and context information.
    • Outputs: Daily predictions, recommendations, and personalized messages.
    • Message Construction: Tone, style, personalization, and reactions to LLM-generated content.
    • Engagement Factors: Ease/difficulty, perceived benefits, and barriers.

Proposed Exit Interview Structure

Introductory Framing

  • Clarify what “the system” means:
    • Inputs provided by participants (EMA, GPS).
    • Predictions and feedback generated by the machine learning model.
    • Messages crafted by the LLM.
  • Emphasize that feedback should focus on algorithmic components, not superficial interface issues (e.g., text message delivery, app design).
  • Acknowledge that current implementation is a first iteration, not the final app experience.

Question Themes

  1. Broad, Open-Ended Questions:
    • Start with general impressions: “Tell me about your experience using the system.”
    • Allow participants to raise unexpected issues before narrowing focus.
  2. System Components:
    • Inputs:
      • What was easy or hard about providing EMA and GPS data?
      • Concerns about invasiveness, time burden, or privacy.
    • Messages:
      • Reactions to message content (risk probability, recommendations, features).
      • Tone and style preferences (e.g., repetitive, engaging, humorous, formal/informal).
      • Understanding that messages were automated and not human-generated.
    • LLM Interaction:
      • Did participants personify the system or feel like they were interacting with an agent?
      • Comfort level with automated feedback.
  3. Costs vs. Benefits:
    • What aspects of the system helped with recovery?
    • What interfered with daily life or felt burdensome?
    • Explore cost-benefit trade-offs (e.g., data sharing vs. perceived value).
  4. Engagement & Adherence:
    • What made it easier or harder to stay engaged?
    • Would participants continue using the system without payment?
    • How did compensation influence participation?
  5. Trust & Confidentiality:
    • Would trust change if the system were offered outside a research context (e.g., as an app)?
    • What would make participants comfortable using it in real-world settings?
  6. Future Use & Recommendations:
    • Would you download and use this app if available in the App Store?
    • Preferred delivery channels (provider referral vs. direct download).
    • Suggestions for improvement (while filtering out irrelevant interface requests).

Design Considerations

  • Funnel Approach:
    • Begin broadly → narrow to specific components → end with future use questions.
  • Probes:
    • Use participant-specific data (e.g., adherence patterns, EMA completion rates) to tailor follow-up questions.
  • Language:
    • Simplify phrasing for clarity (e.g., “What was hard about using the messages?” instead of “What was challenging about engaging with the system?”).
  • Avoid Leading Questions:
    • Frame neutrally; offer examples only if participants struggle to respond.

Additional Recommendations

  • Consider adding Likert-scale items before the interview to guide discussion (e.g., “The system helped me with my recovery” → follow-up: “How did it help?”).
  • Include questions about reading level and message clarity.
  • Explore perceived impact on behavior and recovery.
  • End with a future-oriented question: “Would you use this system if it were available outside of research?”

Logistical Issues

  • Timing:
    • Conduct interviews at exit or dropout, not midpoint (to avoid influencing ongoing engagement).
  • Dropout Bias:
    • Capture feedback from participants who leave early to avoid skew toward completers.
    • Strategies:
      • Offer stepped incentives (e.g., $10 for brief feedback at consent refusal, $50 for full exit interview).
      • Use screening questions for those who decline consent to identify barriers.
  • IRB Considerations:
    • Ensure non-coercive framing when asking reasons for non-participation.
    • De-identify responses; avoid publishing identifiable quotes from single participants.

Next Steps

  • Revise Interview Guide:
    • Incorporate feedback on structure, language, and focus.
    • Add probes for personalized adherence data.
  • Pilot Test:
    • Share draft with qualitative research experts (e.g., Lily, Wisconsin qualitative feedback group).
  • Determine Sample Size & Incentives:
    • Decide whether to interview all participants or a diverse subset.
  • Explore Early Feedback Collection:
    • Develop abbreviated version for non-consent or early dropout cases.

Key Themes

  • Emphasis on algorithmic components (inputs, predictions, LLM messages).
  • Structured exploration of costs, benefits, and barriers.
  • Importance of trust, confidentiality, and real-world adoption factors.
  • Need for inclusive feedback from completers and dropouts.

2.4.13 2025_0624

Preparation for CARDS Advisory Board meetings and planning summer lab meeting topics. CARDS (Community Advisors on Research Design and Strategies) will help ensure equitable technology development and inform recruitment strategies for the Optimize grant.

CARDS Advisory Board Engagement

Purpose

  • Gather community perspectives to:
    1. Improve recruitment and engagement of diverse participants for Optimize.
    2. Inform long-term deployment of the monitoring and support app beyond research settings.
  • Address algorithmic fairness and broader equity concerns by incorporating voices from underserved groups.

Board Composition

  • Two boards from Madison community:
    • Members from community centers, food programs, women’s services, senior groups.
    • Diverse socioeconomic backgrounds; one board predominantly women, another leaning toward low SES and less stable situations.
  • ~12 members per board; meetings capped at ~5 lab attendees to avoid overwhelming participants.
  • Members may or may not have lived experience with SUD (some likely do).

Meeting Logistics

  • Two in-person sessions (July & August), 90 minutes each.
  • Food provided; informal networking encouraged.
  • Flow:
    • Lab introduces purpose and motivation.
    • CARDS facilitates discussion; lab primarily listens.
  • Deliverables:
    • Audio recordings and qualitative analysis report from CARDS.

Discussion Points for CARDS

  • Recruitment Barriers & Comfort:
    • What would make participants comfortable recommending the study/app to friends or family?
    • Barriers to learning about the study and engaging with research.
    • Factors influencing trust in algorithms and recommendations.
  • Technology Use & Burden:
    • Willingness to share geolocation and EMA data.
    • Comfort with push notifications and phone-based participation.
    • Concerns about being “tied to a phone” and device flexibility.
  • Trust & Motivation:
    • How to build trust in research and technology.
    • What information should researchers share about their personal motivation and care for participants.
  • Community Perspectives on SUD:
    • Stigma, cultural/spiritual views, family involvement.
    • Ideas about recovery (ongoing support vs episodic treatment).
    • Sensitive framing to avoid asking boards to speak for entire communities or disclose personal biases.
  • Future Considerations:
    • Automated systems and LLM-generated messaging—how to introduce these concepts later.
    • Potential questions about familiarity with automated tools.

Tasks Assigned

  • Claire:
    • Confirm if CARDS can collect additional participant characterization (insurance, housing, SUD history).
    • Verify whether researchers should share personal motivations during meetings.
  • Susan:
    • Provide compensation details for CARDS to share.
  • Team:
    • Review protocol paper for clarity on participant expectations.

Additional Notes

  • CARDS discouraged asking questions about lived experience or personal biases.
  • Suggested framing: “Imagine a technology app in a research study that collects health-related data and provides feedback—what would make you comfortable engaging or recommending it?”

Summer Lab Meeting Planning

  • Tentative dates: July 15, 22, 29; Aug 5, 19, 26.
  • Goal: Distribute responsibility for leading meetings and selecting topics.
    • Topics can be project-related, new papers, or innovative ideas.
    • Ensure papers suggested are discussion-worthy.
  • Action: Reply to thread with available dates and topic suggestions.
  • Note: Sarah unavailable for last four meetings; one session will review her recommendations.

Key Themes

  • Equity and trust-building in digital health research.
  • Practical barriers to participation and technology adoption.
  • Community-informed strategies for recruitment and engagement.
  • Preparing for qualitative insights to guide Optimize and future app deployment.

2.4.14 2025_0603

Main Objective

Finalize design for Qualtrics panel study to pilot test message tone categories and gather data to inform the Optimize grant. The study aims to:

  • Validate tone categories.
  • Assess user preferences for message styles.
  • Explore fairness/equity implications.
  • Collect actionable insights for personalized messaging in recovery support systems.

Study Components

  1. Tone Preferences Survey:
    • Participants rate their preference for ~7 tone categories (e.g., legitimizing distress, caring/supportive, self-efficacy, acknowledgment of feelings, value affirmation, norms).
    • Use Lyn’s descriptions + one short example per tone dimension.
    • Goal: Identify whether participants have clear tone preferences.
  2. Message Rating Survey:
    • Participants rate messages grouped by risk context blocks (e.g., high risk & increasing, low & stable).
    • Each block starts with an instructional stem: “Imagine this is for a person who is…”.
    • Messages will vary by tone within each context; large item pool ensures variability across participants.
    • Response items (7-point Likert: strongly agree → strongly disagree):
      • “I liked this message.”
      • “I found this message helpful.”
    • Additional exploratory items considered but deprioritized (e.g., “supportive,” “interesting,” “easy to understand”).

Design Considerations

  • Survey Length: Target ~15 minutes to balance data richness and participant burden.
  • Blocking by Context: Easier for assessment but risk of participants rating all messages similarly within a block.
  • Randomization: Each participant sees different message exemplars; prompts remain consistent.
  • Instructional Framing: Emphasize that messages are automated and intended to support engagement with recovery.

Key Research Questions

  1. Tone Validity:
    • Do participants recognize tone categories implicitly?
    • Are ratings more similar within tone categories than across categories?
  2. Preferences & Equity:
    • Do tone preferences differ systematically by demographics (e.g., gender, education)?
    • If so, personalization may reduce disparities in engagement.
  3. Context Interaction:
    • Do preferences vary by risk context (e.g., high/increasing vs. low/stable)?
    • Requires full manipulation of risk contexts within participants.

Tone Categories (Proposed)

  • Legitimization of distress.
  • Caring/supportive.
  • Encouraging/feasibility (self-efficacy).
  • Acknowledgment of feelings.
  • Value affirmation.
  • Norms.
  • Directness (under consideration; may overlap with formal/informal).

Orthogonal Dimensions

  • Formal vs. informal tone (confirmed).
  • Positive tone: All messages positive (avoid extremes).
  • Autonomy: Avoid negative face tone; no explicit autonomy manipulation.
  • Gain/loss framing: Excluded.
  • Message length & complexity: May intersect with tone (formal vs. informal, intellectualizing vs. simple).

Additional Notes

  • Equity Analysis: If tone preferences correlate with demographics, personalization becomes critical.
  • Response Scale: Use agree/disagree anchors for consistency across items.
  • Exploratory Ideas:
    • Compare tone preference ratings vs. actual message ratings.
    • Consider infographic for Optimize (not for panel study).
  • Tasks Assigned:
    • Susan: Summarize decisions for John.
    • CM: Convert summary into prompts for message generation.

Next Steps

  • Finalize tone definitions and examples.
  • Draft tone preference sub-survey and message rating sub-survey.
  • Pre-register analysis plan (including correlation thresholds for collapsing items).
  • Prepare large message pool with tone/context variations.

2.4.15 2025_0506

David Mohr webinar - not specific to optimize, but has some useful engagement suggestions.

Event Overview

  • Hosted by: Northwestern University Center for Behavioral Intervention Technologies (CBITs) and Society for Digital Mental Health (SDMH).
  • Topic: Implementation of digital mental health treatments.
  • Speakers:
    • David Mohr – Director, CBITs; President, SDMH.
    • Trina Histon – Former Kaiser Permanente leader; now at Rōbot.
    • Jenna Carl – Chief Medical Officer, Big Health.

Purpose: Discuss best practices for implementing digital mental health solutions in healthcare systems, based on insights from SDMH’s recently published Digital Mental Health Treatment Playbook.

Key Themes & Insights

1. Implementation Framework

  • Staged Model:
    • PlanningPilot (30–50 patients) → ScalingSustainment.
  • Planning Phase:
    • Stakeholder Mapping: Identify and engage clinicians, operations, IT, legal, contracting early.
    • Understand pain points, care settings, and integration challenges.
    • ABL Principle: “Always Be Learning” – every conversation is an opportunity to understand system needs.
  • Pilot Goals:
    • Clarify purpose (implementation design, not efficacy testing).
    • Define success criteria and post-pilot plan upfront.
    • Avoid “pilot purgatory” where projects stall without scaling.

2. Success Metrics (KPIs)

  • Focus on implementation outcomes, not clinical efficacy:
    • Provider training completion.
    • Provider adoption rates (offers/referrals).
    • Speed to support (time from referral to patient engagement).
    • Patient uptake and activation.
  • Best Practice:
    • Limit KPIs to 5–7 key indicators.
    • Track and share data continuously to identify barriers and optimize processes.
    • Use data + stories to drive engagement and improvement.

3. Provider Engagement & Referral Process

  • What Works:
    • Clinicians personally test tools → builds confidence and credibility.
    • Personalized recommendations tied to patient needs.
    • Digital Navigators: Non-clinical staff who assist with onboarding and troubleshooting.
    • Seamless referral workflows (e.g., via EHR systems like Epic/MyChart).
  • What Doesn’t Work:
    • Vague recommendations without clear follow-up.
    • High-friction processes for patients post-referral.
  • Efficiency Strategies:
    • Provider training upfront; leverage early adopters for peer-to-peer influence.
    • Use orchestration layers (e.g., Zeld, Orca) for streamlined referrals.
    • Nudges/reminders (max two) to encourage patient activation.

4. Patient Engagement & Retention

  • Challenges:
    • Dropout risk similar to psychotherapy; patients face anxiety, low motivation.
  • Strategies:
    • Clear expectations and structured onboarding.
    • Behavioral design principles (nudges, gamification, progress tracking).
    • Notifications: “Nudge, not nag” – allow user control over timing.
    • Show progress visually (mood tracking, milestones).
    • Flexibility: Encourage use during moments of need, not rigid schedules.
  • Trend:
    • Moving from early “text-heavy” apps to consumer-grade UX inspired by platforms like Duolingo.
    • Human support (coaches, navigators) improves engagement but is not always necessary with good design.

5. Integration into Care Pathways

  • Risk: Poor integration → fragmentation, limited impact.
  • Best Practice:
    • Embed digital tools within standard care workflows (primary care, specialty care, collaborative care).
    • Use tools to augment therapy (homework, between-session reinforcement).
    • Offer as first-line or adjunctive treatment for conditions like insomnia, anxiety, depression.
  • Key Consideration:
    • Match tool to care setting and patient readiness.
    • Digital tools can normalize and prepare patients for therapy if initially resistant.

6. Reimbursement Landscape

  • New CMS G-Codes (Jan 2025):
    • Cover device supply and provider onboarding.
    • Additional codes for treatment management (20-min increments, mostly asynchronous).
  • Requirements:
    • FDA clearance under classification for computerized behavioral therapy for psychiatric disorders.
    • Clinical evidence of safety and efficacy.
  • Current State:
    • Medicare covers codes; commercial payer adoption is emerging.
    • Pricing variability across Medicare Administrative Contractors (MACs); national pricing expected later.
  • Challenges:
    • Low initial reimbursement rates (e.g., $129) may threaten sustainability.
    • Need broader payer coverage and fair pricing to support scalability.
  • Future Outlook:
    • Codes create a pathway for integration into routine care.
    • Industry feedback and data will shape pricing and coverage expansion.

7. Prevention & Early Intervention

  • Interest from payers and employers in preventive digital solutions.
  • Barriers:
    • High regulatory bar for prevention claims.
    • Unclear reimbursement models for non-treatment interventions.
  • Opportunities:
    • Early education and engagement tools with demonstrated preventive benefits.

Actionable Takeaways

  • For Health Systems:
    • Engage stakeholders early; clarify goals and success metrics.
    • Train providers and leverage digital navigators.
    • Integrate tools into workflows; minimize friction for referrals.
  • For Developers:
    • Prioritize UX, personalization, and behavioral design.
    • Prepare for regulatory compliance and reimbursement pathways.
  • For Policy Makers:
    • Expand coverage and pricing models to support sustainability.
    • Encourage interoperability and integration standards.

2.4.16 2025_0428

Purpose of Meeting

  • Review the concept and measurement of Recovery Capital and its potential integration into the Optimize grant project.
  • Discuss psychometric properties of the Recovery Capital measure and implications for outcome sensitivity.
  • Explore opportunities to enhance predictive models by incorporating monthly survey data, contextual features, and expanded EMA items.
  • Address feature engineering strategies and next steps for model development.

Key Highlights

1. Recovery Capital Measure

  • Construct Overview:
    • Recovery Capital refers to resources that support recovery (social, physical, human, cultural).
    • Historically under-measured; recent efforts aim to capture both positive and negative ends of the construct.
  • Measure Development:
    • Iterative process involving stakeholder input, item refinement, and psychometric evaluation.
    • Reliability assessed via internal consistency (Cronbach’s alpha) and test-retest stability.
      • Subscale alphas: generally strong (0.81), one lower (0.65).
      • Test-retest conducted over one week; weighted kappa used for item-level stability.
    • Validation: limited to correlations with related constructs; could be more robust.
  • Strengths:
    • Comprehensive approach to item development.
    • Incorporates protective factors, not just risk.
  • Limitations:
    • MTurk used for pilot testing—concerns about data quality and participant authenticity.
    • Construct changes slowly; may not show rapid shifts during short interventions.

2. Integration into Optimize Project

  • Planned Use:
    • Include Recovery Capital as a secondary outcome alongside heavy drinking days and drinking days.
    • Rationale:
      • Potentially more sensitive than distal clinical outcomes.
      • Captures intermediate changes in recovery resources.
  • Measurement Strategy:
    • Baseline and 4-month follow-up as primary comparison points.
    • Two-month data may serve as proxy for missing four-month data.
  • Recruitment Considerations:
    • Target participants in early recovery (1–3 months abstinent).
    • Avoid extremely early stages to reduce dropout risk.

3. Enhancing Predictive Models

  • Monthly Survey Data:
    • Includes scales for craving, self-efficacy, stress, quality of life, social support, and living arrangements.
    • Plan:
      • Feature engineer most recent score for each scale/subscale.
      • Avoid adding change scores or medians due to missingness and limited predictive gain.
      • Use subscales where available for dimensionality reduction.
  • Contextual Features:
    • GPS and metadata can approximate recovery capital indicators:
      • Social Support: Time spent with supportive vs. non-supportive contacts.
      • Housing Stability: Home address changes, overnight location patterns.
      • Access to Resources: Visits to healthcare or recovery-related locations.
      • Pleasant Alcohol-Free Activities: Locations rated as enjoyable and low-risk.
    • Potential to combine GPS patterns with context surveys for richer features.
  • Audio Data Opportunity:
    • Daily voice updates (transcribed and stored) could provide predictive signals via:
      • Language embeddings for themes (stress, motivation).
      • Acoustic analysis for affective states.
    • Compliance challenges noted; future use may involve weekly recordings or improved UX.

4. EMA Expansion for Optimize

  • Current EMA design was risk-focused; lacked protective factors.
  • Proposed additions:
    • Affect: Simplify to one item (predominant emotion) or retain discrete emotions.
    • Sleep Quality: Reintroduce after sensor discontinuation.
    • Positive Behaviors: Checklist for recovery-supportive actions (exercise, hydration, coping strategies).
    • Anticipated Pleasant Events: Add forward-looking items.
    • Physical Health & Pain: Broader wellness indicators.
    • Guilt/Shame: Capture internalizing emotional states relevant to SUD.
  • Goal: Balance comprehensiveness with low participant burden (1x daily administration).

5. Action Items & Next Steps

  • Feature Engineering:
    • Madison & Susan to:
      • Review monthly survey scoring.
      • Implement feature extraction for most recent scale scores.
      • Organize features by category for integration with EMA and GPS data.
  • EMA Revision:
    • Team to propose additional items for risk and protective factors.
    • Compare Risk 1 vs. Risk 2 EMA sets; incorporate lessons learned.
  • Model Update Plan:
    • Combine Optimize and Risk 1 datasets for retraining.
    • Add new features from monthly surveys and context data.
  • Future Exploration:
    • Investigate feasibility of leveraging audio transcripts for predictive modeling.
    • Consider LLM-based categorization and embeddings for open-ended responses.

Additional Notes

  • Madison announced new role at WPS (Medicare/Medicaid fraud analytics) starting May 20.
  • Lily departing for graduate studies in Copenhagen (Master’s in Cognition & Communication).
  • Team acknowledged need for privacy considerations in future app development, especially for SMS and audio data.

2.4.17 2025_0421

Purpose

  • Review progress on psychoeducation pages and intervention modules for personalized recommendations.
  • Discuss structure, tone, and content for modules linked to top predictive features.
  • Explore design considerations for usability, engagement, and scalability.

Context

  • Participants receive daily messages based on predictive features (e.g., craving, past use, future efficacy, risky locations).
  • First exposure to a feature → psychoeducation page (background, normalization, terminology).
  • Subsequent exposures → short, actionable interventions (1–5 minutes).
  • Goal: Ensure interventions are clear, feasible, and engaging without overwhelming participants.

Key Topics Discussed

1. Psychoeducation Pages

  • Purpose:
    • Normalize experiences (e.g., cravings, lapses).
    • Explain why the feature matters for recovery.
    • Prepare participants for recommended activities.
  • Design:
    • Introductory paragraph + key points in bold.
    • Reading level: 6th grade target for accessibility.
    • Keep short and focused; avoid embedding full activities.
  • Debate:
    • Should psychoeducation include a reflection prompt (e.g., “How does this apply to you?”)?
      • Consensus: Yes, brief reflection is helpful.
    • Avoid mixing psychoeducation with full interventions to prevent perceived burden.

2. Intervention Modules

  • Structure:
    • Each activity page starts with:
      • Link back to relevant psychoeducation.
      • Brief explanation of why the activity matters.
      • Clear, step-by-step instructions.
    • Include visuals or diagrams to enhance comprehension.
  • Examples for Craving:
    • Urge Surfing:
      • Analogy: Craving as a wave (rises quickly, peaks, subsides).
      • Needs unpacking for participants unfamiliar with mindfulness.
      • Ideal format: Intro text + illustrative image + steps.
    • Other activities:
      • 5-4-3-2-1 grounding (sensory distraction).
      • Delay and distraction techniques.
      • Visualization exercises.
      • “Cold water reset” (physiological calming).
      • Recovery reminders and reframing strategies.
  • Feedback:
    • Short text-only versions feel too minimal; prefer rich format with context and visuals.
    • Use existing resources (e.g., Matrix manual) where copyright allows; avoid reinventing content.

3. Media and Interactivity

  • Current approach: Static web pages (Qualtrics link → HTML/QMD pages).
  • Options discussed:
    • Include guided audio/video for activities (e.g., mindfulness, urge surfing).
    • Provide downloadable PDFs or encourage bookmarking.
  • Challenges:
    • Limited interactivity without a dedicated app.
    • Risk of participants accessing all content at once (reduces graded exposure).
  • Proposed solution:
    • Maintain sequential delivery.
    • Consider curated “My Activities” list:
      • Updated daily with links to previously recommended activities.
      • Accessible via message link or EMA prompt.
      • Grouped by category with brief descriptions.

4. Personalization and Feedback

  • Idea: Allow participants to thumbs-up preferred activities for future tailoring.
  • Barriers:
    • Requires bidirectional data flow and database integration.
    • Feasible in future app-based versions; complex for current web-based setup.
  • Alternative:
    • Sequential exposure to variations (e.g., text → video → advanced version).
    • Build curated history for participant reference.

5. Motivation vs Confidence (Future Efficacy Feature)

  • EMA question: “How confident are you in staying abstinent over the next week?”
  • Discussion:
    • Motivation and confidence are distinct constructs:
      • Motivation = desire to pursue goal.
      • Confidence = belief in ability to achieve goal.
    • Current design combines them; may need separate psychoeducation and activities.
  • Proposed approach:
    • Psychoeducation page introduces both concepts broadly.
    • Activities tailored:
      • Motivation: Reflect on values, reasons for change.
      • Confidence: Identify skills, plan incremental steps.
    • Examples:
      • “Three reasons” exercise (write down why abstinence matters).
      • Scaling and boosting (rate confidence/motivation, plan to increase by one point).

Design Principles Agreed

  • Keep psychoeducation concise; avoid overwhelming participants.
  • Activities should be actionable in 1–5 minutes.
  • Visuals are essential for clarity and engagement.
  • Sequential delivery to prevent content overload.
  • Use existing validated resources where possible (Matrix manual, clinical tools).
  • Plan for future scalability (personalization, interactivity) but prioritize feasibility now.

Next Steps

  • Sarah:
    • Finalize psychoeducation pages for all top predictors.
    • Expand activity descriptions with context and visuals.
    • Check copyright for external resources; share with Lily for verification.
  • Team:
    • Explore technical feasibility of curated activity lists and optional video integration.
    • Review additional modules (past use, risky locations) for structure and clarity.
  • Madison:
    • Continue model updates; integrate new predictors (e.g., post-lapse commitment).
  • Deadline: Initial module set ready by May 1 for implementation testing.

2.4.18 2025_0303

Purpose

  • Define structure and content for modules recommended to participants based on predicted risk features.
  • Clarify how psychoeducation and activities will be organized and delivered.
  • Discuss integration of EMA and GPS predictors into recommendation logic.

Context

  • Participants receive daily messages with varying components:
    1. Lapse risk probability (high/low).
    2. Trend information (increasing/decreasing/stable).
    3. Top predictive feature (e.g., craving).
    4. Personalized recommendation (activity/support).
  • Recommendations may appear with or without other message components.
  • Goal: Make recommendations feel tailored and clinically meaningful.

Key Design Decisions

Module Structure

  • Organize modules around predictors, not therapeutic class:
    • Example: If craving is top predictor → recommend craving-related activities.
    • Rationale: Easier to explain relevance and personalize interventions.
  • Two layers of psychoeducation:
    1. Risk factor level (e.g., craving, past use, stress).
    2. Intervention framework level (e.g., mindfulness for urge surfing).
  • Activities should be standalone, with optional links to deeper psychoeducation.

Delivery Flow

  • First exposure to a predictor:
    • Link to psychoeducation page explaining why the feature matters.
  • Subsequent exposures:
    • Direct link to a specific activity.
    • Activity page includes link back to psychoeducation for review.
  • Avoid overwhelming participants with all activities at once.

Example: Craving Module

  • Psychoeducation essentials:
    • Cravings are normal, temporary, predictable, and controllable.
    • Understanding these principles primes participants for interventions.
  • Sample activities:
    • Urge surfing: Visualize craving as a wave that peaks and subsides.
    • Grounding distraction (5-4-3-2-1): Engage senses to redirect attention.
    • Reflective exercise: Identify physical sensations and thoughts during craving.
  • Design considerations:
    • Reading level: Target ~5th grade for accessibility.
    • Formatting: Use tables, callouts, and inclusive language.

Other Predictors & Proposed Psychoeducation

  1. Past Use:
    • Normalize lapses as common and non-moral failures.
    • Emphasize learning from lapses and planning for future situations.
    • Activities: Relapse planning, mindfulness reflection, thought records.
  2. Abstinence Goal:
    • Normalize dips in motivation.
    • Encourage reflection on values and reasons for abstinence.
    • Activities: Motivation-building exercises.
  3. Risky Situations:
    • Past vs future:
      • Past: Reflect on triggers and consequences.
      • Future: Plan for anticipated risks.
    • Consider merging into one page with content on avoidable vs unavoidable risks.
  4. Stressful Events:
    • Similar approach to risky situations.
    • Normalize stress, teach coping strategies, and planning.
    • May combine past/future stressors into one module for simplicity.
  5. Well-being Module:
    • For low-risk days or protective features.
    • Reinforce positive behaviors and coping strategies.

Integration of GPS Features

  • Top GPS predictors (from prior analysis):
    • Time spent at self-defined risky locations.
    • Type of location (e.g., bars, restaurants).
    • Time spent at places marked as “avoid in recovery.”
  • GPS features likely map to risky situations module but may require:
    • Location-specific suggestions (e.g., route changes, alternative activities).
    • Psychoeducation on why tracking location matters.

Personalization Considerations

  • Potential enhancements:
    • Include recent EMA ratings in messages (e.g., “You rated craving 7/10”).
    • Acknowledge past vs future stressors in message text.
  • Challenges:
    • Adds complexity (branching logic, hard-coded conditions).
    • Decision deferred until module framework is stable.

Technical & Implementation Notes

  • Website for modules:
    • Prefer Quarto (QMD) for easy updates and formatting.
    • Include imagery, tables, and clear navigation.
  • Message flow:
    • First link → psychoeducation page.
    • Subsequent links → activity page with back-link to psychoeducation.
  • Avoid embedding evaluation questions in message flow; consider EMA follow-up instead.

Outstanding Questions

  • Should past/future stressors and risky situations share one psychoeducation page?
  • How to handle GPS-driven recommendations when risk factor is ambiguous?
  • To what extent should personalization (e.g., EMA scores in messages) be implemented now vs future iterations?

Next Steps

  • Sarah:
    • Draft psychoeducation pages for all predictor categories.
    • Outline activity pools for each category.
  • Madison:
    • Analyze EMA data for abstinence goal variability (yes/no/uncertain counts).
    • Post summary to Slack.
  • Team:
    • Review GPS feature integration once combined EMA+GPS models are available.
    • Plan visual design for module website (Susan to assist).
  • Future discussion:
    • Explore personalization strategies and branching logic feasibility.

2.4.19 2025_0224

Purpose

  • Finalize message design manipulations and primary outcome measures for the optimized grant project.
  • Discuss engagement as the primary outcome and its relationship to message content.
  • Explore additional language features and measurement timing.

Core Focus

  • Goal: Determine if message manipulations matter for engagement and eventual clinical benefit.
  • Clarification:
    • The study is not evaluating the entire system or app UX.
    • It is focused on message content manipulations as the first component of the system.

Message Manipulations

Four Primary Components (Factorial Design)

  1. Daily lapse probability (risk level for that day).
  2. Trend information (whether lapse probability is increasing, decreasing, or stable).
  3. Top risk feature contributing to lapse probability.
  4. Personalized daily activity/support recommendation.
  • Fully factorial design → 16 conditions (from none to all components).
  • Between-subjects manipulation (each participant stays in one condition).

Additional Language Features (Not Primary Manipulations)

  • Caring/supportive language:
    • Likely included in all messages for variability and engagement.
    • Adds human tone beyond formulaic content.
  • Acknowledging feelings:
    • Potentially personalized using EMA data (affect, stressors, pleasant events).
    • Could be manipulated message-to-message or always on.
    • Adds complexity (requires LLM integration and personalization logic).

Study Design

  • Burn-in period: 1 week for EMA and geolocation data collection.
  • Intervention period: 16 weeks (4 months).
  • Measurement points:
    • Baseline.
    • Midpoint (8 weeks).
    • Endpoint (16 weeks).
    • Possible additional measure at 4 weeks.
  • Daily messages delivered via Qualtrics + Twilio; EMA in evening; optional link to support website.

Measurement Strategy

Baseline Measures

  • Demographics (inclusive categories for participant identity).
  • Alcohol use history and AUD symptom count (self-report SKID or equivalent).
  • WHO ASSIST for broad substance use.
  • Comorbidity screening (broadband psychiatric measures).
  • Propensity to trust algorithms (baseline trust measure).

Primary Outcome

  • Engagement:
    • Defined as click-through to Qualtrics message page (binary: read/not read).
    • Count of days engaged per 8-week block (analyzed as Poisson distribution).
    • Intent-to-treat analysis:
      • No missing data (dropouts scored as zero engagement).
      • Avoids bias from complete-case or imputation methods.
  • Rationale:
    • Engagement is closest proxy to message usefulness.
    • Clinical outcomes (e.g., drinking days) considered too distal for primary measure.

Secondary/Exploratory Outcomes

  • Clinical outcomes (heavy drinking days, drinking days).
  • EMA completion (exploratory, not primary).
  • Potential message-level ratings (trust/usefulness) via EMA (evening):
    • Pros: Captures temporal granularity.
    • Cons: Missing data if EMA skipped.

Key Discussion Points

  • Engagement as Proxy:
    • Engagement chosen because clinical outcomes are premature for this phase.
    • Goal: Evaluate message components, not maximize engagement artificially.
  • Avoid Ceiling Effects:
    • Adding strong engagement drivers (e.g., streaks, gamification) could mask differences between message conditions.
    • Supportive language may increase engagement but is not clinically impactful alone.
  • Personalization Complexity:
    • Using EMA affect data for personalization adds technical and design challenges.
    • May be deferred to future studies or CM’s Hilldale project.
  • Covariates:
    • Baseline trust and phone usage patterns could help explain engagement variability.
  • Message-Level Questions:
    • Avoid embedding in message flow (complicates navigation).
    • Prefer EMA-based follow-up questions (e.g., “Was today’s message useful?”).
  • Technical Considerations:
    • Qualtrics click-through provides reliable engagement measure.
    • Duration on page is unreliable (confounded by idle time).
    • Website-based tracking adds complexity without clear benefit.

Future Considerations

  • Explore moderators:
    • Mean lapse probability and variability across study may influence usefulness of risk feedback.
  • Consider personalization strategies:
    • Affect-based language.
    • Progress tracking (e.g., streaks, milestones).
  • Balance innovation vs complexity:
    • Advanced personalization may be deferred to future grants.

Action Items

  • Confirm engagement (click-through) as primary outcome.
  • Finalize baseline measures (demographics, AUD severity, trust).
  • Decide on inclusion of message-level EMA questions.
  • Determine feasibility of affect-based personalization vs simplified caring language.
  • Document rationale for avoiding streaks/gamification in this phase.

2.4.20 2025_0127

Purpose of Meeting

  • Refine primary outcome measures for the optimized grant project focused on messaging interventions for SUD recovery.
  • Discuss trust, usefulness, and engagement as key constructs.
  • Review potential measurement approaches and feedback from collaborator Rachel (expert in personalized messaging for depression).

Project Context

  • Goal: Develop and test a monitoring and support system for substance use disorder (SUD) recovery.
  • Intervention: Factorial manipulation of message content (4 components):
    1. Daily lapse probability.
    2. Change in lapse probability over 2 weeks.
    3. Key risk feature contributing to lapse risk.
    4. Personalized support recommendation.
  • Participants: 4–6 months in study; messages delivered via Qualtrics + Twilio, geolocation via Follow Me app, and links to support modules on a standalone site.
  • Current system = test bed, not final app (limited UX focus).

Key Outcome Domains

  1. Trust in system/messages.
  2. Usefulness of messages (distinct from app usability).
  3. Engagement (continued interaction with system).
  4. Clinical outcomes (heavy drinking days, drinking days) – secondary.

Measurement Strategy

1. Baseline

  • Individual differences in propensity to trust automated systems (covariate).
  • Potential items from trust literature.

2. Midpoint & Endpoint

  • Trust and usefulness measures to assess impact of message content over time.
  • Candidate measures:
    • TOAST (Trust in Automated Systems Test) – recent, covers multiple trust dimensions.
      • Pros: Clear language, short, multidimensional.
      • Cons: Some items focus on system mechanics vs perceived accuracy.
    • HCTM – alternative trust measure.
  • Rachel suggested:
    • Digital Working Alliance Inventory (bond, task relevance).
    • Subjective Engagement Scale (affective, cognitive, behavioral).
    • Custom usability measures if needed.

3. Daily Level

  • Engagement:
    • Click-through rate to messages (proxy for engagement).
    • % of days clicked through over study period.
  • Optional quick questions after each message:
    • Trust in message (accuracy/relevance).
    • Usefulness for recovery.
    • Keep short (1–2 items) to avoid burden.

4. Rachel’s Contribution

  • Shared homegrown measure from depression messaging project:
    • 7 items assessing message qualities (supportive, relevant, easy to understand, interesting, helpful).
    • Developed for engagement-focused intervention using reinforcement learning.
    • Preliminary factor analysis: Mostly unidimensional construct.
  • Suggested:
    • Consider custom measure for trust/usefulness tailored to SUD context.
    • Clarify what “trust” means (accuracy, personalization, timing).
    • Possibly adapt provider trust items for system trust.

Key Discussion Points

  • Trust vs Usefulness:
    • Trust = belief in accuracy of information.
    • Usefulness = actionable/helpful for recovery.
    • Both likely necessary for sustained engagement.
  • Granularity:
    • Daily ratings capture interaction with specific message content.
    • Midpoint/endpoint ratings capture overall perception.
  • Challenges:
    • Trust is multidimensional (algorithm accuracy, data handling, timing).
    • Risk of undermining trust if algorithm outputs incorrect info (e.g., misinterpreted geolocation).
  • Measurement Design:
    • Avoid overly broad heterogeneous scales.
    • Focus on constructs linked to engagement and manipulated message components.
    • Consider adding items for negative reactions (e.g., stress from high-risk feedback).
  • Future Use:
    • Data will inform which message components drive usefulness and engagement.
    • Insights guide next-phase app design and personalization strategies.

Action Items

  • Team members to post 1–2 key takeaways in Slack (Optimized channel).
  • Continue refining:
    • Daily vs periodic measurement balance.
    • Item wording for trust/usefulness.
    • Integration with EMAs for retrospective message evaluation.
  • Explore Rachel’s measure and potential adaptation for SUD context.
  • Plan next meeting to finalize measurement framework.

2.4.21 2024_1209

This was a presentation practice talk about optimize.

Key Objectives

  • Develop risk monitoring systems integrated into mobile health apps.
  • Predict lapse events (distinct from relapse) as early warning signs.
  • Use machine learning models to deliver personalized, timely interventions.

Background

  • Alcohol and opioid use disorders are chronic conditions requiring lifelong management.
  • Recovery is dynamic and complex; initial treatment alone is insufficient.
  • Continuous monitoring and accessible, affordable solutions are needed.
  • NIH-supported project under UW Addiction Research Center aims to:
    • Address barriers to accessibility.
    • Build temporally precise risk monitoring systems for smartphones.

Model Development

1. OUD Model

  • Sample: 336 individuals, early to mid recovery, 1–12 months medication-assisted treatment.
  • Data: Daily EMA (Ecological Momentary Assessment) self-reports on mood, cravings, substance use.
  • Algorithm: XGBoost (outperformed Random Forest, k-NN, Elastic Net).
  • Validation: 5×10 cross-validation; AUROC used for imbalanced data.
  • Performance:
    • Median AUROC = 0.94 (excellent).
    • High fairness across demographic groups (minimal AUROC differences).
  • Top Features:
    1. Past opioid use.
    2. Confidence in recovery.
    3. Urge to use.
  • Feature Insights:
    • Past use strongly increases lapse risk.
    • Confidence shows nonlinear effect (low confidence ↑ risk; very high confidence ↓ risk).
    • Urge shows opposite nonlinear trend.

2. AUD Model

  • Sample: 151 individuals, moderate to severe AUD, early recovery.
  • Added Features: Geolocation (GPS) for contextual awareness.
  • Performance:
    • AUROC = 0.85 (slightly lower due to 24-hour prediction windows vs hourly).
  • Important Features:
    • Past drinking days.
    • Future efficacy (likelihood of drinking next week).
    • Risky location (self-reported + GPS).
    • Alcohol availability, location type, location valence.
  • Insights:
    • Past use and future efficacy dominate predictions.
    • Risky location impact varies across individuals (mostly neutral overall).
    • Location valence (emotional tone) can increase risk for some participants.

Personalization & Shapley Analysis

  • Shapley values used to:
    • Identify daily risk/protective features per participant.
    • Show variability in feature importance over time.
  • Examples:
    • Participant A: Past use ↑ risk around day 35; risky location coincides.
    • Participant B: Past use consistently protective; location valence ↑ risk.
  • Implication: Personalized recommendations are essential as risk factors differ across individuals and time.

Implementation Plans

  • Deliver tailored recommendations based on current risk features.
  • Explore language, tone, and format to minimize burden.
  • Develop measures for trust, engagement, and usability.
  • Goal: Engagement with support messages acts as a protective factor.

2.4.22 2024_1202

Meeting Goal

  • Understand the types of interventions/messages to deliver in treatment messages.
  • Identify how these messages differ to inform feature selection for message generation.
  • Review examples from existing apps and resources.

Relapse Prevention Model (RPM)

Focus: Skill-based microinterventions that are:

  • Complete in themselves (not requiring ordered learning).
  • Contextual and personalizable (more than conventional CBT).
  • Brief and actionable (3–5 minutes).

Five Categories of Microinterventions

  1. Building Awareness
    • Awareness of drinking habits, feelings, cravings.
    • Often used early in treatment; less relevant for advanced stages.
  2. Coping Skills
    • Largest category for us.
    • Brief interventions for distress/cravings (e.g., urge surfing, meditation, distraction).
    • Alternative actions to substance use.
  3. Planning
    • Skills for reflection and safety planning after events.
    • Less variation; tied to specific events.
  4. Enhance Motivation
    • Boost confidence and motivation to abstain.
    • Address overconfidence or low motivation after setbacks.
  5. Cognitive Restructuring
    • Most complex; drawn from CBT.
    • Changes thinking patterns to prevent hazardous situations.
    • High cognitive load; likely for participants doing well.
    • Requires explanation and prior exposure.

Matrix Considerations:

  • Cognitive restructuring likely for low-risk or post-lapse scenarios.
  • Avoid expanding the matrix further (too high-dimensional).

Message Selection & Variability

  • Select most important feature of the day or rotate among features above threshold.
  • Add variability in:
    • Message content (empathy, feasibility).
    • Presentation style.
  • Even within one exercise, add features for variation.

Sources & Examples

  1. Sarah’s Treatment Group Worksheets
    • Focus on Building Awareness & Coping Skills.
    • Could link to worksheets or break into smaller messages.
    • Open use, no copyright.
  2. ASIM Addiction Monitoring Manual
    • Worksheets for monitoring progress.
    • Consider breaking down for variability but maintain usefulness.
  3. Other Apps
    • I AM SOBER: High variability, but language can feel pejorative.
    • WiseMind: Visually appealing interface.
    • SmartRecovery / RecoveryPath: Interactive worksheets, minimal typing.
    • Center for Healthy Minds: Uses timers for reflection.

Design Considerations

  • Activity Size: Shortest helpful unit (3–5 minutes).
  • Interface Consistency: All messages should have a uniform look.
  • Visual Appeal:
    • Explore budget for design improvements.
    • Consult Mary Anderson (PR) for recommendations.
  • Engagement Measurement:
    • Track interaction with complex activities (e.g., “reflect on your last lapse”).
    • Interactive formats preferred over PDFs.

Additional Notes

  • Avoid overwhelming users with all content at once.
  • Consider foundational work for low-risk participants:
    • Protective factors.
    • Motivational enhancement (e.g., listing positive things).
  • Plan for interactive, guided experiences rather than static worksheets.

2.4.23 2024_1125

Project Status Updates

  • Model development shows promising results:
    • 1x daily EMA performs as well as 4x daily EMA.
  • Current model performance:
    • 0.85, down from 0.90 → requires investigation.
  • GPS features integrated without impacting performance.
  • Team investigating:
    • Impact of 24-hour vs hour-by-hour prediction windows.

Next Steps

  • Determine feature importance within the model (global and local).
  • Develop methodology for identifying daily important features for individual users.
  • CM to present updated app development progress (planned for early spring semester).
  • Finalize timing strategy for daily EMA collection and feedback messages.

Implementation Plans

  • System will run daily risk prediction models overnight.
  • App will collect:
    • 1x daily EMA
    • GPS data
  • One support message sent daily (likely in morning).
  • EMA collection planned for late afternoon/early evening.
  • Team to discuss optimal timing windows for user interactions.

Action Items & Assignments

  • Susan & PI:
    • Experiment with Notion for meeting transcripts and organization.
  • Madison:
    • Continue development of risk prediction models.
    • Investigate drop in model performance (0.90 → 0.85).
    • Analyze feature importance globally.
  • CM:
    • Present updated app development progress (early spring semester).
  • All Grad Students:
    • Restart usage of Asana for project management.
  • Still Needs Assignment:
    • Develop methodology for identifying daily important features for individual users.
    • Determine optimal timing windows for user interactions.
    • Finalize strategy for daily EMA collection and feedback messages.

2.4.24 2024_1118

1. Pivot in DTx Approach

  • JC’s ideas have shifted from AS treatment to post-treatment support.
  • Focus on individuals who:
    • Start with intensive treatment (ER detox, PHP, inpatient).
    • Begin with standard outpatient therapy (e.g., CBT programs, 10-week or 12-session).
    • Engage in continuing care or long-term therapy for skill reinforcement.
  • New direction:
    Everyone should receive skills-based training (outpatient, PHP, or inpatient).
  • Sarah’s suggestion:
    DTx could run concurrently with initial treatment to maintain engagement after treatment ends (similar to VA PTSD apps).

2. Treatment Goals

  • Abstinence-based vs alternative goals:
    • Harm reduction.
    • Mastery of skills presented in treatment (e.g., completing 12-week CBT).
    • For severe cases, abstinence remains appropriate.
  • Skills focus:
    • Coping skills review (2–3 min).
    • Coping skills practice (2–10 min).
    • Not teaching new skills—reinforcing existing ones.

3. Role of EMA and Engagement

  • EMA as intervention:
    • Encourages daily self-monitoring.
    • Reflection likely reduces relapse risk.
    • Feedback improves over time → better compliance.
  • App engagement:
    • Keeps users involved.
    • Builds autonomous self-monitoring habits.

4. Skill Delivery Considerations

  • Skills don’t need psych jargon—plain language works.
  • Challenge:
    • App = one-directional teaching.
    • Some skills require bidirectional interaction (e.g., tele-mental health before app use if no prior treatment).
  • Support modules:
    • Provide enough content for assumptions.
    • Avoid positioning app as a substitute for initial treatment (to prevent defunding of therapy).

5. Recruitment

  • Partner with local clinics.
  • JC has interest from 2 treatment providers; more may follow.
  • Intake process:
    • Ask what tools patients learned.
    • Map those to app toolkit or review during intake.
  • Timeline:
    • Build modules in spring semester.
    • Launch in ~6 months.

6. Second Issue: Model Limitations

  • Current models trained only on R1 & R2 inputs.
  • Original design aimed for passive system → EMA underdeveloped.
  • Narrow recommendation system:
    • Missing features JC wishes were included.
    • Limited ability to change now (predictive validity concerns).
  • Opportunity:
    • Collect better data during this project.
    • Use for future algorithms with broader feature sets.

7. Next Steps

  • Discuss with treatment providers:
    • What’s missing in current models?
    • Valuable for future improvements (time constraints prevented discussion today).

2.4.25 2024_0923

1. Core Questions

  • How does the model recommend actions today?
  • How do we map model output to the clinical model matrix?
  • What are we validating?
    → Whether the 4 information bits provided are correct and useful for clinical outcomes and engagement (proposed).
    → Engagement should be disconnected from a specific therapeutic so it can apply to any DTx.

2. Proposed Outcomes

  • Clinical Outcomes (exploratory):
    • Drinking days (3m and 6m)
    • Heavy drinking days
  • Engagement (definition pending)
  • Trust:
    • Fundamental question: Do people trust the algorithm’s information?
    • Trust was not part of Sub1 but requested by reviewers.
    • Errors in recommendations could erode trust.

3. Design Considerations

  • Factorial Design:
    • 16 cells from 0–4 components (2^4 factorial, between-subjects).
    • Allows aggregation over time.
  • OPT Design:
    • Requires a single primary outcome for decision-making.
    • Clinical outcomes likely not suitable as primary (premature).
      • Recommendations may improve over time.
      • Outcomes depend on external factors (e.g., quality of DTx).
    • Trust may be better as primary OPT outcome.
      • Engagement is downstream of trust.

4. Trust Measurement

  • Grant proposed Korber’s Trust in Automation Scale.
    • John thinks it’s clunky.
  • Plan:
    • Review existing trust measures in next meeting.
    • Likely develop a custom scale based on:
      • Personal characteristics
      • History with model
      • Baseline trait-like differences in trust.
  • Risks:
    • Geolocation info → perceived as invasive.
    • EMA features → less risky (self-reported).
    • Lapse probability → risky (iatrogenic effects, self-fulfilling prophecy).

5. Additional Measures

  • At 3/6 months:
    • Usability scales from app research (e.g., usefulness, benefit).
    • Even if not app-based, some items may apply.
  • Daily quick measures:
    • Embedded in feedback messages (via Qualtrics):
      • Single-item Likert on trust
      • Single-item Likert on usefulness
    • Rich for follow-up analyses (patterns within individuals).
    • Need to finalize wording.
  • Timing:
    • Feedback vs EMA message timing is critical (pin for later discussion).

6. Engagement Metrics

Two simplified approaches:

  1. System Engagement:
    • Click-through/read rate of support messages.
    • Response to trust/usefulness questions.
  2. Recovery Engagement:
    • EMA item: “Did you do anything yesterday to support your recovery?” (Y/N)
      • If Yes → optional detail (checkboxes or free text).
      • Avoid incentivizing “No” to skip follow-up.

7. Other Notes

  • Consider populating TLFB with daily reports and allow corrections.
  • All messages should include a statement about feasibility for recovery.
  • Future discussion: timing of feedback vs EMA.

2.5 CARDS Meeting Notes

We contracted with CARDS for 4 meetings to have their community volunteer groups provide feedback on our materials.

Each of the following opens the final report provided by CARDS from the S: drive