For example, a photocopier can automatically sort and collate copies.
sentence describing examples of a concept
For example, a photocopier can automatically sort and collate copies.
sentence describing examples of a concept
Ithasbeenusedtodescribeindividuals,groups,andcommunitiesusingcomputers.
Wehaveusedittodiscussvariousapplications,fromausertypingonasmart-phonetoateamofinformationworkerscommunicatingviaemail.
P8 said: "as I get familiar with this system [C3], I feel more skilled" to use the highlighted and grayed phrases.
any sentence that describes a user's emotional (positive or negative) response to any condition in the experiments.
It is worth noting that P3 and P8 mentioned feeling more comfortable with the more familiar visualization in C1 and C2 during their first impression of the conditions.
any sentence that describes a user's emotional (positive or negative) response to any condition in the experiments.
The inclusion of counterfactuals often resulted in a substantial increase in precision, indicating that the models were better able to correctly classify relevant instances while reducing false positives.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
By visualizing these consistent pattern rules, users may be better understanding the behavior of the model through inference projection.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
Mocha addresses two seemingly contradictory objectives: (1) generating labeled data that diversifies the training dataset to aid the model's learning, and (2) maintaining structural consistency across the batches of data presented to users to support their cognitive processes.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
The results of our study indicate that participants spent significantly less time annotating batches of counterfactuals when they were rendered according to SAT compared to other conditions i.e., supporting the participants' selective focus on the varying phrases, rather than phrases that stay consistent.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
From a cognitive perspective, the theme color aligns with the human's (theorized) structural mapping engine [27] by making relational discrepancies between the original and counterfactual examples more explicit.
return any single sentence that describes an explicit or implicit connection to theory
Estes and Hasson [17] highlights the significant role of bringing salience to 'non-alignable' differences.
return any single sentence that describes an explicit or implicit connection to theory
The last two prior works also combine Variation Theory (VT) and SAT together, as we did (i.e., a corollary of SAT referred to as Analogical Transfer/Learning Theory).
return any single sentence that describes an explicit or implicit connection to theory
Estes and Hasson [17] highlights the significant role of bringing salience to "non-alignable" differences.
return any single sentence that describes an explicit or implicit connection to theory
Estes and Hasson [17] argue that while alignable differences can be more straightforward and easier for comparison, non-alignable differences can also provide key information that might otherwise remain overlooked.
return any single sentence that describes an explicit or implicit connection to theory
By incorporating theories such as Structural Alignment Theory and Variation Theory, it aims to support the learning of both the human and the model.
return any single sentence that describes an explicit or implicit connection to theory
This symbiotic relationship stems from the fact that Structural Alignment Theory (SAT) enhances the salience of differences, while the way we used Variation Theory (VT) to generate contradicting examples across the boundaries of labels ensures that these differences are conceptually informative.
return any single sentence that describes an explicit or implicit connection to theory
It states that understanding and sensemaking involve mapping the relationships between elements, especially in complex and ambiguous tasks.
return any single sentence that describes an explicit or implicit connection to theory
Structural Alignment Theory states that humans naturally look for structural mapping between representations of objects to help them understand, compare, and infer relationships between said objects.
return any single sentence that describes an explicit or implicit connection to theory
According to Variation Theory, learners better understand concepts by observing variations along critical features (dimensions of variation) that define that concept and, separately, observing variations along superficial features that do not define that concept—all while other features, when possible, are held constant.
return any single sentence that describes an explicit or implicit connection to theory
Mocha exemplified the application of human cognition and concept learning theories in the interactive machine learning pipeline to support the negotiation of conceptual boundaries for bi-directional human-AI alignment.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
This pattern of selective attention suggests that the visual cues provided by Mocha effectively guided participants to focus on more relevant information within the context of unchanged text when making their labeling decisions.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
Overall, the incorporation of counterfactuals has generally improved the models' F1 scores, driven largely by the improvements in precision. This suggests that counterfactuals have effectively improved performance without necessitating a significant trade-off between precision and recall.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
The inclusion of counterfactuals often resulted in a substantial increase in precision, indicating that the models were better able to correctly classify relevant instances while reducing false positives. This improvement suggests that the counterfactuals provided essential information that helped refine the models' decision boundaries.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
By visualizing these consistent pattern rules, users may be better understanding the behavior of the model through inference projection [26]. This can not only boosts the model's performance but also enable participants to validate or correct the model during the interactive training process.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
Thus, the integration of both theories enables users to efficiently process and compare variations, leading to more informed decisions and a clearer understanding of the model's behavior.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
By helping users see alignable differences, SAT-based rendering helps users focus on key variations that are essential to changing the data item's label, making it easier to interpret the effects of changes and their significance.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
We argue that these two theories form a symbiotic relationship (Fig. 6). Variation Theory provides the conceptual basis for generating structurally consistent differences, while Structural Alignment Theory (SAT) enhances the user's ability in recognizing and processing these differences.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
Participants were able to efficiently focus on key differences between the original and counterfactual examples, which facilitated more efficient annotations.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
The results from our user study suggest that both the participants and the model benefited from the Variation Theory (VT)-based counterfactuals and Structural Alignment Theory (SAT)-based rendering.
statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.
Thus, the integration of both the-ories enables users to efficiently process and compare variations,leading to more informed decisions and a clearer understanding ofthe model’s behavior.
Specifically, they underscore the need for co-adaptive systems that can evolve along with users' mental models and definitions of labels.
any sentence that describes explicit design implications
This requires the development of interfaces and visualizations that demystify the generated data, allowing systematic variation and coverage across the concept space.
any sentence that describes explicit design implications
By visualizing these consistent pattern rules, users may be better understanding the behavior of the model through inference projection [26].
any sentence that describes explicit design implications
By incorporating theories such as Structural Alignment Theory and Variation Theory, it aims to support the learning of both the human and the model.
any sentence that describes explicit design implications
Such studies could determine whether these non-alignable comparisons enhance user performance and elicit deeper insights in human-AI collaborative systems.
any sentence that describes explicit design implications
Most previous research in counterfactual generation has focused on the model side by either generating counterfactuals to improve the model's performance or explaining its behaviors post hoc.
any single sentence that compares and contrasts this work with prior work.
Variation Theory provides the conceptual basis for generating structurally consistent differences, while Structural Alignment Theory (SAT) enhances the user's ability in recognizing and processing these differences.
return any single sentence that describes an explicit or implicit connection to theory
While SAT-based rendering supported human sensemaking in both Gero et al. [29] and Mocha, we also show that the combination of VT and SAT support the model's learning.
any single sentence that compares and contrasts this work with prior work.
This finding is consistent with previous work that supports users' sense-making of text, e.g., by modulating text saliency. Specifically, Gu et al. [32] and Gero et al. [29] both found improved reading efficiency and comprehension with saliency-modulating text renderings.
any single sentence that compares and contrasts this work with prior work.
In decision making, SAT argues that people tend to focus on alignable differences—features that can be directly compared—rather than on differences that cannot be easily aligned.
return any single sentence that describes an explicit or implicit connection to theory
Structural Alignment Theory (SAT) [27] is a cognitive theory that explains how people make sense of concepts by comparing relational structures between two items.
return any single sentence that describes an explicit or implicit connection to theory
Specifically, we use Variation Theory of learning [44] which states that for learning to occur, some aspects that define the concept being learned must vary while others are held constant.
return any single sentence that describes an explicit or implicit connection to theory
According to SAT, humans compare two similar entities by trying to find structural alignments between them, and then comparing corresponding elements, with a special focus on differing aligned elements.
return any single sentence that describes an explicit or implicit connection to theory
VT posits that human learning occurs when learners experience variation across critical and superficial aspects of a concept—through exposure to contrasting examples that systematically vary along different critical and superficial feature dimensions.
return any single sentence that describes an explicit or implicit connection to theory
To analyze the annotation efficiency, we first conducted a Kruskal-Wallis rank sum test [39] to determine if there were statistically significant differences in annotation time across the three conditions, because our data violated the homogeneity of variances assumption, making non-parametric methods more appropriate.
return any single sentence that describes data analysis done on data collected by the authors when running human subjects experiments.
Mohamed et al. (2020) put forth the idea of dismantling power assymmetries to resist data colonialism.
sentence that refers to a theory
Couldry and Mejias (2019) propose 'data colonialism' as a new form of colonialism to make sense of the use of large amounts of data by a small group of corporate and government actors.
sentence that refers to a theory
Taken together, these findings almost unanimously show that, on average, AI-supported writing decreases but does not eliminate writer's feelings of ownership, underscoring the need for a larger theory of AI participation in the creative process.
sentence that refers to a theory
This can be understood through the frame of precarious work [5]; as writers feel that their work is increasingly precarious, the power differential between themselves and the organizations seeking to train LLMs grows larger.
sentence that refers to a theory
information scholars argue that data should be collected and used in accordance with data subjects' perspectives on the acceptable use of their data [29, 40, 83].
sentence that refers to a theory
using SemanticCommit, we recorded all instances of edits, check for conflicts, make change, local, and global resolution actions using telemetry.
sentences describing methods the authors used; one sentence at a time
A sub-task was considered a failure if the participant was unable to complete it within the time limit.
sentences describing methods the authors used; one sentence at a time
Finally, we conducted an informal interview about their experience.
sentences describing methods the authors used; one sentence at a time
After both tasks were completed, participants filled out a final survey to compare the two conditions.
sentences describing methods the authors used; one sentence at a time
We chose GPT-4o for performance and latency reasons, as it performed optimally against our evals.
sentences describing methods the authors used; one sentence at a time
We also ran evaluations of model latency and classification performance under varying false positive rates for the following LLMs by OpenAI: GPT-4o, GPT-4o-mini, and o3-mini.
sentences describing methods the authors used; one sentence at a time
For each task, participants were tasked with integrating three new pieces of information into the memory, one at a time ("sub-tasks").
sentences describing methods the authors used; one sentence at a time
We ensured each list was 30 items long as our pilot studies suggested this was long enough that manual detection starts to become unwieldy (users need to scroll up and down the document), but short enough that participants could become familiar in a short period.
sentences describing methods the authors used; one sentence at a time
We adapted two intent specifications from our evals: Mars Game Design Document and Financial Advice AI Agent Memory, as these tasks mapped to the two paradigmatic types covered in Sections 2 and 2.1 (design documents, and AI memory of the user).
sentences describing methods the authors used; one sentence at a time
We recruited 12 participants (7 female, 5 male) through the mailing lists of two research universities and one multinational technology company.
sentences describing methods the authors used; one sentence at a time
We chose OpenAI's ChatGPT Canvas as a baseline for five reasons: (i) it is a popular, commercially available tool, hence it is likely familiar to users; (ii) it provides a document editing view, where users can select text and ask GPT to rewrite it, or chat with an AI to make global edits; (iii) it employs a similar class of model (GPT-4o); (iv) it supports similar editing features as SemanticCommit like inline text selection, conflict highlighting, and a diff view, while adding free-form editing; and (v) similar interfaces like Anthropic Artifacts tended to rewrite the specification entirely, and did not offer Canvas's "diff" view to allow for a fair comparison.
sentences describing methods the authors used; one sentence at a time
With participant consent, we recorded audio and screen-casts, and participants were encouraged to think aloud.
sentences describing methods the authors used; one sentence at a time
Four coauthors created the evals, and two coauthors manually double-checked all conflicts, a process that took several days.
sentences describing methods the authors used; one sentence at a time
We ran one pilot study with five users of our card-based interface, and a second with four users of a revised interface.
sentences describing methods the authors used; one sentence at a time
Our explorations went through substantial iterations and prompt prototyping over a period of eight months, evolving in response to two pilot studies and progressing from a card-based interface to a list of texts.
sentences describing methods the authors used; one sentence at a time
We iterated on prompts using ChainForge [5] by setting up an evaluation pipeline against our datasets, which allowed us to observe the effects of prompt changes and model choices.
sentences describing methods the authors used; one sentence at a time
To measure statistical significance, we used Mann–Whitney–Wilcoxon tests and report the p-values.
sentences describing methods the authors used; one sentence at a time
For qualitative analysis, the first author performed open coding on participant responses and audio transcripts to identify themes, which were used to interpret the qualitative results.
sentences describing methods the authors used; one sentence at a time
In the post-task surveys, we collected self-reported NASA Task Load Index (TLX) scores, Likert-scale ratings for ease of use, and responses on how well the AI helped participants identify, understand, and resolve semantic conflicts.
sentences describing methods the authors used; one sentence at a time
Each condition had a time limit of 15 minutes, after which the participant completed a post-task survey.
sentences describing methods the authors used; one sentence at a time
Before each task, participants received a tutorial on the assigned tool and were given five minutes to explore it using a test document.
sentences describing methods the authors used; one sentence at a time
Both the order of task assignment and tool assignment were counterbalanced and randomly assigned.
sentences describing methods the authors used; one sentence at a time
We conducted a controlled within-subjects study with mixed methods, comparing SemanticCommit with a baseline interface.
sentences describing methods the authors used; one sentence at a time
We run end-to-end on our four eval datasets using GPT-4o and GPT-4o-mini and report the mean ± stddev for accuracy, precision, recall, and F1 scores for the three approaches in Figure 5.
sentences describing methods the authors used; one sentence at a time
We compare our end-to-end system against two simpler methods: (i) DropAllDocs, which adds all documents to the context for conflict classification; and (ii) InkSync [56] which generates a JSON list of string-replace operations.
sentences describing methods the authors used; one sentence at a time
In order to minimize relevance assessment issues, we apply a PageRank-based relevance ranking over the KG, akin to HippoRAG [36].
sentences describing methods the authors used; one sentence at a time
We implement the back-end using a knowledge-graph (KG) RAG architecture [36] consisting of two phases: pre-processing and inference.
sentences describing methods the authors used; one sentence at a time
Through a within-subjects study with 12 participants comparing SemanticCommit to a chat-with-document baseline (OpenAI Canvas), we find differences in workflow: half of our participants adopted a workflow of impact analysis when using SemanticCommit, where they would first flag conflicts without AI revisions then resolve conflicts locally, despite having access to a global revision feature.
sentences describing methods the authors used; one sentence at a time
We implement the back-end using a knowledge-graph (KG) RAG architecture [36] consisting of two phases: pre-processing and inference.
sentences describing methods the authors used; one sentence at a time
In order to minimize relevance assessment issues, we apply a PageRank-based relevance ranking over the KG, akin to HippoRAG [36].
sentences describing methods the authors used; one sentence at a time
We compare our end-to-end system against two simpler methods: (i) DropAllDocs, which adds all documents to the context for conflict classification; and (ii) InkSync [56] which generates a JSON list of string-replace operations.
sentences describing methods the authors used; one sentence at a time
We run end-to-end on our four eval datasets using GPT-4o and GPT-4o-mini and report the mean ± stddev for accuracy, precision, recall, and F1 scores for the three approaches in Figure 5.
sentences describing methods the authors used; one sentence at a time
We conducted a controlled within-subjects study with mixed methods, comparing SemanticCommit with a baseline interface.
sentences describing methods the authors used; one sentence at a time
Both the order of task assignment and tool assignment were counterbalanced and randomly assigned.
sentences describing methods the authors used; one sentence at a time
Before each task, participants received a tutorial on the assigned tool and were given five minutes to explore it using a test document.
sentences describing methods the authors used; one sentence at a time
Each condition had a time limit of 15 minutes, after which the participant completed a post-task survey.
sentences describing methods the authors used; one sentence at a time
In the post-task surveys, we collected self-reported NASA Task Load Index (TLX) scores, Likert-scale ratings for ease of use, and responses on how well the AI helped participants identify, understand, and resolve semantic conflicts.
sentences describing methods the authors used; one sentence at a time
For qualitative analysis, the first author performed open coding on participant responses and audio transcripts to identify themes, which were used to interpret the qualitative results.
sentences describing methods the authors used; one sentence at a time
To measure statistical significance, we used Mann–Whitney–Wilcoxon tests and report the p-values.
sentences describing methods the authors used; one sentence at a time
We iterated on prompts using ChainForge [5] by setting up an evaluation pipeline against our datasets, which allowed us to observe the effects of prompt changes and model choices.
sentences describing methods the authors used; one sentence at a time
Our explorations went through substantial iterations and prompt prototyping over a period of eight months, evolving in response to two pilot studies and progressing from a card-based interface to a list of texts.
sentences describing methods the authors used; one sentence at a time
We ran one pilot study with five users of our card-based interface, and a second with four users of a revised interface.
sentences describing methods the authors used; one sentence at a time
Four coauthors created the evals, and two coauthors manually double-checked all conflicts, a process that took several days.
sentences describing methods the authors used; one sentence at a time
With participant consent, we recorded audio and screen-casts, and participants were encouraged to think aloud.
sentences describing methods the authors used; one sentence at a time
These semantic conflicts require dedicated support to detect, visualize, and resolve. Semantic conflict resolution interfaces must go beyond visualizing what changes were made, to what changes could be made, where they should be made, and what the effects might be. This resembles feedforward: affordances that help the user foresee the impact of an action [67, 93].
sentences describing connections to theory; one sentence at a time
Today with LLMs, we are less limited by this constraint, and solutions to the problem of human-machine communication might be better found in cybernetics theory [9] than static formalism.
sentences describing connections to theory; one sentence at a time
This reflects the principle of feedforward [67, 93] in communication theory—"a needed prescription or plan for a feedback, to which the actual feedback may or may not confirm" [79]—where a communicator provides "the context of what one was planning to talk about" [64, p. 179-80] in order to "pre-test the impact of [its output]" on the listener [34, p. 65].
sentences describing connections to theory; one sentence at a time