The difficulties inherent in the evaluation of educational software are described in terms of the tradeoffs between internal, external, and ecological validity. Larger issues in evaluation research design and computer-based instruction are highlighted by primary and metaanalytic studies designed to reveal the effects of computer simulations in psychology classrooms and laboratories. The effectiveness of classroom and laboratory computer activities depends on how the inclusion of software, as well as the evaluation process itself, changes the entire instructional process.