Friday, September 20, 2013

National Science Foundation funds research that puts engineering design processes under a big data "microscope"

The National Science Foundation has awarded us $1.5 million to advance big data research on engineering design. In collaboration with Professors Şenay Purzer and Robin Adams at Purdue University, we will conduct a large-scale study involving over 3,000 students in Indiana and Massachusetts in the next five years.

This research will be based on our Energy3D CAD software that can automatically collect large process data behind the scenes while students are working on their designs. Fine-grained CAD logs possess all four characteristics of big data defined by IBM:
  1. High volume: Students can generate a large amount of process data in a complex open-ended engineering design project that involves many building blocks and variables; 
  2. High velocity: The data can be collected, processed, and visualized in real time to provide students and teachers with rapid feedback; 
  3. High variety: The data encompass any type of information provided by a rich CAD system such as all learner actions, events, components, properties, parameters, simulation data, and analysis results; 
  4. High veracity: The data must be accurate and comprehensive to ensure fair and trustworthy assessments of student performance.
These big data provide a powerful "microscope" that can reveal direct, measurable evidence of learning with extremely high resolution and at a statistically significant scale. Automation will make this research approach highly cost-effective and scalable. Automatic process analytics will also pave the road for building adaptive and predictive software systems for teaching and learning engineering design. Such systems, if successful, could become useful assistants to K-12 science teachers.

Why is big data needed in educational research and assessment? Because we all want students to learn more deeply and deep learning generates big data.

In the context of K-12 science education, engineering design is a complex cognitive process in which students learn and apply science concepts to solve open-ended problems with constraints to meet specified criteria. The complexity, open-endedness, and length of an engineering design process often create a large quantity of learner data that makes learning difficult to discern using traditional assessment methods. Engineering design assessment thus requires big data analytics that can track and analyze student learning trajectories over a significant period of time.
Deep learning generates big data.

This differs from research that does not require sophisticated computation to understand the data. For example, in typical pre/post-tests using multiple-choice assessment, the selection data of individual students are directly used as performance indices -- there is basically no depth in these self-evident data. I call this kind of data usage "data picking" -- analyzing them is just like picking up apples already fallen to the ground (as opposed to data mining that requires some computational efforts).

Process data, on the other hand, contain a lot of details that may be opaque to researchers at first glance. In the raw form, they often appear to be stochastic. But any seasoned teacher can tell you that they are able to judge learning by carefully watching how students solve problems. So here is the challenge: How can computer-based assessment accomplish what experienced teachers (human intelligence plus disciplinary knowledge plus some patience) can do based on observation data? This is the thesis of computational process analytics, an emerging subject that we are spearheading to transform educational research and assessment using computation. Thanks to NSF, we are now able to advance this subject.

Sunday, September 15, 2013

Measuring the effects of an intervention using computational process analytics

"At its core, scientific inquiry is the same in all fields. Scientific research, whether in education, physics, anthropology, molecular biology, or economics, is a continual process of rigorous reasoning supported by a dynamic interplay among methods, theories, and findings. It builds understanding in the form of models or theories that can be tested."  —— Scientific Research in Education, National Research Council, 2002
Actions caused by the intervention
Computational process analytics (CPA) is a research method that we are developing in the spirit of the above quote from the National Research Council report. It is a whole class of data mining methods for quantitatively studying the learning dynamics in complex scientific inquiry or engineering design projects that are digitally implemented. CPA views performance assessment as detecting signals from the noisy background often present in large learner datasets due to many uncontrollable and unpredictable factors in classrooms. It borrows many computational techniques from engineering fields such as signal processing and pattern recognition. Some of these analytics can be considered as the computational counterparts of traditional assessment methods based on student articulation, classroom observation, or video analysis.

Actions unaffected by the intervention
Computational process analytics has wide applications in education assessments. High-quality assessments of deep learning hold a critical key to improving learning and teaching. Their strategic importance has been highlighted in President Obama’s remarks in March 2009: “I am calling on our nation’s Governors and state education chiefs to develop standards and assessments that don’t simply measure whether students can fill in a bubble on a test, but whether they possess 21st century skills like problem-solving and critical thinking, entrepreneurship, and creativity.” However, the kinds of assessments the President wished for often require careful human scoring that is far more expensive to administer than multiple-choice tests. Computer-based assessments, which rely on the learning software to automatically collect and sift learner data through unobtrusive logging, are viewed as a promising solution to assessing increasingly prevalent digital learning.

While there have been a lot of work on computer-based assessments for STEM education, one foundational question has rarely been explored: How sensitive can the logged learner data be to instructions?

Actions caused by the intervention.
According to the assessment guru Popham, there are two main categories of evidence for determining the instructional sensitivity of an assessment tool: judgmental evidence and empirical evidence. Computer logs provide empirical evidence based on user data recording—the logs themselves provide empirical data for assessment and their differentials before and after instructions provide empirical data for evaluating the instructional sensitivity. Like any other assessment tools, computer logs must be instructionally sensitive if they are to provide reliable data sources for gauging student learning under intervention. 


Actions unaffected by the intervention.
Earlier studies have used CAD logs to capture the designer’s operational knowledge and reasoning processes. Those studies were not designed to understand the learning dynamics occurring within a CAD system and, therefore, did not need to assess students’ acquisition and application of knowledge and skills through CAD activities. Different from them, we are studying the instructional sensitivity of CAD logs, which describes how students react to interventions with CAD actions. Although interventions can be either carried out by human (such as teacher instruction or group discussion) or generated by the computer (such as adaptive feedback or intelligent tutoring), we have focused on human interventions in this phase of our research. Studying the instructional sensitivity to human interventions will enlighten the development of effective computer-generated interventions for teaching engineering design in the future (which is another reason, besides cost effectiveness, why research on automatic assessment using learning software logs is so promising).

The study of instructional effects on design behavior and performance is particularly important, viewing from the perspective of teaching science through engineering design, a practice now mandated by the newly established Next Generation Science Standards of the United States. A problem commonly observed in K-12 engineering projects, however, is that students often reduce engineering design challenges to construction or craft activities that may not truly involve the application of science. This suggests that other driving forces acting
Distribution of intervention effect across 65 students.
on learners, such as hunches and desires for how the design artifacts should look, may overwhelm the effects of instructions on how to use science in design work. Hence, the research on the sensitivity of design behavior to science instruction requires careful analyses using innovative data analytics such as CPA to detect the changes, however slight they might be. The insights obtained from studying this instructional sensitivity may result in the actionable knowledge for developing effective instructions that can reproduce or amplify those changes.

Our preliminary CPA results have shown that CAD logs created using our Energy3D CAD tool are instructionally sensitive. The first four figures embedded in this post show two pairs of opposite cases with one type of action sensitive to an instruction that occurred outside the CAD tool and the other not. This is because the instruction was related to one type of action and had nothing to do with the other type. The last figure shows that the distribution of instructional sensitivity across 65 students. In this figure, the largest number means higher instructional sensitivity. A number close to one means that the instruction has no effect. From the graph, you can see that the three types of actions that are not related to the instruction fluctuate around one whereas the fourth type of action is strongly sensitive to the instruction.

These results demonstrate that software logs can not only record what students do with the software but also capture the effects of what happen outside the software.

Wednesday, August 28, 2013

Modeling the hydrophobic effect of a polymer

There are many concepts in biochemistry that are not as simple as they appear to be. These are things that tend to confuse you if you mull over them. Over the years, I have found osmosis such a thing. Another such thing is hydrophobicity. (As a physicist, I love these puzzles!)

Figure 1: More "polar" solvent on the right.
In our NSF-funded Constructive Chemistry project with Bowling Green State University, Prof. Andrew Torelli and I have identified that the hydrophobic effect may be one of the concepts that would benefit the most from a constructionism approach, which requires students to think more deeply as they must construct a sequence of simulations that explain the origin of this elusive effect. Most students can tell you that hydrophobicity is "water-hating" as their textbooks simply have so written. But this layman's term itself is not accurate and might lend itself to a misconception as if there existed some kind of repulsive force between a solute molecule and the solvent molecules that makes them "hate" each other. An explanation of the hydrophobic effect involves quite a few fundamental concepts such as intermolecular potential and entropy that are cornerstones of chemistry. We would like to see if students can develop a deeper and more coherent understanding while challenged to use these concepts to create an explanatory simulation using our Molecular Workbench software.

Andrew and I spent a couple of weeks doing research and designing simulations to figure out how to make such a complex modeling challenge realistic for his biochemistry students to do. This blog post summarizes our initial findings.

Figure 2. The radii of gyration of the two polymers.
First we decided that we would like to set this challenge on the stage of protein folding. There are few problems in biochemistry that are more fundamental than protein folding. So this would be a good brain teaser that could stimulate student interest. But protein folding is such a complex problem. So we would like to start with a simple 2D polymer that is made of identical monomers. This polymer is just a chain of Lennard-Jones particles linked by elastic bonds. The repulsion core of the Lennard-Jones potential models the excluded volume of each monomer and the elastic bonds link them together as a chain. There is no force that maintains the angles of the chain. So the particles can rotate freely. This model is very rough, but it is already an order of magnitude better than the ideal chain, which assumes a polymer as a random walk and neglects any kind of interactions among monomers.

Figure 3. Identical solvents (weakly polar).
Next we need a solvent model. For simplicity, each solvent molecule is represented by a Lennard-Jones particle. Again, this is a very rough model for water as solvent as it neglects the angular dependence of hydrogen bonds among water molecules. A better 2D model for water is the Mercedes-Benz model, so called because its three-arm model for hydrogen bonding resembles the Mercedes-Benz logo. We will probably include this hydrogen bonding model in our simulation engine in the future, but for now, the angular effect may be secondary for the purpose of this modeling project.

As with themselves, the polymer and solvent molecules interact with each other through a Lennard-Jones potential. Now, the question is: Are these interactions we have in hands sufficient to model the hydrophobic effect? In other words, can the nature of hydrophobicity be explained by using this simple picture of interactions? Would Occam's razor be good in this case? I feel that this is a crucial key to our Constructive Chemistry project: If a knowledge system can be reduced to only a handful of rules students can learn, master, and apply in a short time without being too frustrated, the chance of succeeding in guiding them towards learning through construction-based inquiry and discovery would be much higher. Think about all those successful products out there: LEGO, Minecraft, Algodoo, and so on. Many of them share a striking similarity: They are all based on a set of simple building blocks and rules that even young children can quickly learn and use to construct meaningful objects. Yet, from the simplicity rises extremely complex systems and phenomena. We want to learn from their tremendous successes and invent the overdue equivalents for chemistry and biology. The Constructive Chemistry project should pave the road for that vision.
Figure 4. Identical solvents (strongly polar).

Back to modeling the hydrophobic effect: Does our simple-minded model work? To answer this question, we must be able to investigate the effect of each factor. To do so, we set up two compartments separated by a barrier in the middle. Then we put a 24-bead polymer chain into one of them and then copy it to another. In order for them not to move to the edges or corners of the simulation box (if they stay near the edges then they are not fully solvated), we pin their centers down using an elastic constraint. Next we will put different types of solvent particles into the two compartments. We also use some scripts to keep the temperatures on both sides identical all the time and export the radii of gyration of the two polymers to a graph. The radius of gyration of a polymer approximately describes its dimension.

By keeping everything else but one factor identical in the two compartments, we can investigate exactly what is responsible for the hydrophobic effect for the polymers (or its relative importance). Our hypothesis at this point is that the hydrophobic effect would be more pronounced if the solvent-solvent interaction is stronger. To test this, we set the Lennard-Jones attraction between solvent B (right) particles to be three times stronger than that between solvent A particles, while keeping everything else such as mass and size exactly the same. Figure 1 shows a series of snapshots taken from a nanosecond-long simulation (this model has 550 particles in total, but on my Lenovo X230 tablet it runs speedily). The results show that the polymer on the right folds into a hairpin-like conformation with its two freely-moving terminals pointing outwards from the solvent, suggesting that it attempts to leave the solvent (but cannot because it is pinned down). And this conformation and location last for a long time (in fact most of the time during the simulated nanosecond). In comparison, the polymer on the left has no stable conformation or location -- it is randomly stretched in the solvent most of the time and does not prefer any specific location. I think this is the evidence for the hydrophobic effect in two senses: 1) The polymer attempts to separate from the solvent; and 2) the polymer curls up to make room for more contacts among the solvent particles (this is related to the so-called hydrophobic collapse in the study of protein folding). The second can be further visualized by comparing the radii of gyration (Figure 2), which consistently differ by 2-3 angstroms.

Note that we did not introduce any special interaction between the polymers and the solvent particles of either type. The interaction between the polymer with a solvent particle is exactly the same in both compartments. The only difference is the solvent-solvent interaction. The difference in the simulation results for the two polymers is all because it is energetically more favorable for the solvent particles in the right compartment to stay closer. After numerous collisions (this is sometimes called entropy-driven), the hairpin conformation emerges as the winner for the polymer on the right.
Figure 5: Higher temperatures.

To make sure that there is no mistake, we ran another simulation in which the two solvents were set to be identically weak-polar. Figure 3 shows that there was no clear formation of a stable conformation for either polymer in a nanosecond-long simulation. Neither polymer curled up.

Next we set the two solvents to be identically strong-polar. Figure 4 shows that the two polymers both ended up in a hairpin conformation in a nanosecond-long simulation.

Another test is to raise the temperature but keep the solvent-solvent interaction in the right compartment three times stronger than that in the left compartment. Can the polymer on the right keep its hairpin conformation when heated? Negative, as shown in Figure 5. This actually is related to denaturation, a process in which a protein loses its stable conformation due to heat (or other external stimuli).

These simulations suggest that our simple-minded model might be able to explain the hydrophobic effect and allow students to explore a variety of variables and concepts that are of fundamental importance in biochemistry. Our next steps are to transfer the modeling work we have done to something students can also do. To accomplish this goal, we will have to figure out how to scaffold the modeling steps to provide some guidance.

Wednesday, August 14, 2013

Some thoughts and variations of the Gas Frame (a natural user interface for learning gas laws)

A natural user interface (NUI) is the user interface that is based on natural elements or natural actions. Interacting with computer software through a NUI simulates everyday experiences (such as swiping a finger across a touch screen to move a photo in display or just "asking" a computer to do something through voice commands). Because of this resemblance, a NUI is intuitive to use and requires little or no time to learn. NUIs such as touch screen and speech recognition have become commonplace on new computers.

As the sensing capability of computers becomes more powerful and versatile, new types of NUI emerge. The last three years have witnessed the birth and growth of sophisticated 3D motion sensors such as Microsoft Kinect and Leap Motion. These infrared-based sensors are capable of detecting the user's body language within a physical space near a computer with varied degrees of resolution. The rest is how to use the data to create meaningful interactions between the user and a certain piece of computer software.

Think about how STEM education can benefit from this wave of technological innovations. Being scientists, we are especially interested in how these capabilities can be leveraged to improve learning experiences in science education. Thirty years of development, mostly funded by federal agencies such as the National Science Foundation, have produced a wealth of virtual laboratories (aka computational models or simulations) that are currently being used by millions of students. These virtual labs, however, are often criticized for not being physically relevant and not providing hands-on experiences commonly viewed as necessary in practicing science. We now have an opportunity to partially remedy these problems by connecting virtual labs to physical realities through NUIs.

What would a future NUI for a science simulation look like? For example, if you teach physical sciences, you may have seen many versions of gas simulations that allow students to interact with them through some kind of graphical user interface (GUI). What would a NUI for interacting with a gas simulation look like? How would that transform learning? Our Gas Frame provides an example of implementation that may give you something concrete to think about.

Figure 1: The Gas Frame (the default configuration).
In the default implementation (Figure 1), the Gas Frame uses three different kinds of "props" as the natural elements to control three independent variables related to a gas: A warm or cold object to heat or cool the gas, a spring to exert force on a piston that contains the gas, and a syringe to add or remove gas molecules. The reason that I call these objects "props" is because, like in film making, they mostly serve as close simulations to the real things without necessarily performing the real functions (you don't want a prop gun to shoot real bullets, do you?).

The motions of the gas molecules are simulated using a molecular dynamics method and visualized on the computer screen. The volume of the gas is calculated in real time using the molecular dynamics method based on the three physical inputs. In addition to the physical controls through the three props, a set of virtual controls are available on the screen for students to interact with the simulation such as viewing the trajectory path or the kinetic energy of a molecule. These virtual controls support interactions that are impossible in reality (no, we cannot see the trajectory of a single molecule in the air).

The three props can control the gas simulation because a temperature sensor, a force sensor, and a gas pressure sensor are used to detect student interactions with them, respectively. The data from the sensors are then translated into inputs to the gas simulation, creating a virtual response to a real action (e.g., molecules are added or subtracted when the student pushes or pulls a syringe) and a molecular interpretation of the action (e.g., molecules run faster or slower when temperature increases or decreases).

Like in almost all NUIs, the sensors and the data they collect are hidden from students, meaning that students do not need to know that there are sensors involved in their interactions with the gas simulation and they do not need to see the raw data. This is unlike many other activities in which sensors play a central role in inquiry and must be explicitly explained to students (and the data they collected must be visually presented to students, too). There are definitely advantages of using sensors as inquiry tools to teach students how to collect and analyze data. Sometimes we even go extra miles to ask students to use a computer model to make sense of the data (like the simulation fitting idea I blogged before). But that is not the reason why the National Science Foundation funded innovators like us to do.

The NUIs for science simulations that we have developed in our NSF project all use sensors that have been widely used in schools, such as those from Vernier Software and Technology. This makes it possible for teachers to reuse existing sensors to run these NUI apps. This decision to build our NUI technology on existing probeware is essential for our NUI apps to run in a large number of classrooms in the future.

Figure 2: Variation I.
Considering that not all schools have all the types of sensors needed to run the basic version of the Gas Frame app, we have also developed a number of variations that use only one type of sensor in each app.

Figure 2 shows a variation that uses two temperature sensors, each connected to the temperature of the virtual gas in a compartment. The two compartments are separated by a movable piston in the middle. Increasing or decreasing the temperature of the gas in the left or right compartment through heating or cooling the thermal contacts in which the sensors are applied will cause the virtual piston to move accordingly, allowing students to explore the relationships among pressure, temperature, and volume through two thermal interactions in the real world.

Figure 3: Variation II.
Figure 3 shows another variation that uses two gas pressure sensors, each connected to the number of molecules of the virtual gas in a compartment through an attached syringe. Like in Variation I, the two compartment are separated by a movable piston in the middle. Pushing or pulling the real syringes will cause molecules to be added or removed from the virtual compartments, allowing students to explore the relationships among number of molecules, pressure, and volume through two tactile interactions.

If you don't have that many sensors, don't worry -- both variations will still work if only one sensor is available.

I hear you asking: All these sounds fun, but so what? Will students learn more from these? If not, why bother to go through these extra troubles, compared with using an existing GUI version that needs nothing but a computer? I have to confess that I cannot answer this question at this moment. But in the next blog post, I will try to explain our plan for figuring this out.