Trustworthy AI-Assisted Software Engineering
Ongoing study of measurement validity, residual failure after passing tests, persistent unsafe artifacts from coding agents, complementarity of assurance gates, and whether empirical conclusions survive model generations.
-
Evaluation Protocol Validity
Whether an evaluation protocol measures agent capability at all, and whether protocol choice alone can reverse a published conclusion. Examined through repeated runs and protocol transfer.
- Multi-Run Evaluation
- Protocol Transfer
- Reproduction
-
Residual Failures Behind Green Tests
How much behavioral failure remains in AI-generated patches that pass every test, and under what conditions an acceptance test can count as reliability evidence.
- Behavioral Divergence
- Lower-Bound Estimation
- Reliability Evidence
-
Persistent Unsafe Artifacts
Whether repository-level prompt injection ends at attack success or leaves durable unsafe artifacts in the codebase, and whether those artifacts pass an ordinary development pipeline.
- Repository Injection
- Artifact Persistence
- Pipeline Escape
-
Complementarity of Assurance Gates
Whether tests, static analysis, and CI checks detect different defects or the same ones, and which combinations justify their cost inside an assurance pipeline.
- Conditional Detection
- Gate Complementarity
- Cost–Benefit
-
Longitudinal Reproduction
Whether a conclusion obtained today survives a change of model generation, and how model drift bounds the useful lifetime of an empirical finding.
- Model Drift
- Replication
- Temporal Validity