In this study, we computationally analyzed published structures of Hepatitis B virus core protein (HBc) bound with different core protein assembly modulators (CAMs) that interfere with HBV assembly. We focused on comparing the difference in CAM binding between two mutations of HBc, one that forms flat sheets, and one that forms icosahedra similar to the wild type virus capsid. We did this by aligning all currently published HBc-CAM structures by their CAM binding pockets. This allowed us to both quantitatively and qualitatively compare the CAM binding pockets of these two different structure types. We find that there are critical differences in the angle of interaction, capsid orientation angle, CAM pocket shapes and sizes, and CAM interacting residues between icosahedral and sheet based HBc structures.
Title:
Assembly-active and -inactive forms of HBV capsid protein provide distinctly different binding sites for capsid assembly modulators
This dataset contains confocal imaging and single-cell data used to characterize the developmental trajectories of Arabidopsis lateral and median nectaries. The repository includes:
*3D cell tracking: 3D segmentation and quantitative imaging tracking to characterize the developmental patterns of lateral and median nectaries. The data highlights differences in cell expansion, proliferation, and volumetric growth of the nectary epidermis and nectary sub-epidermal cells/parenchyma cells.
*Single-Cell transcriptomics: Gene expression and marker gene analysis identifying nectary-specific cell clusters.
Title:
Live imaging confocal data, marker genes and nectary cell trajectory datasets.
The Concordant Crop Sequence Boundaries improves the stabilty and accuracy of the USDA Crop Sequence Boundaries by altering the agglomeration method to be based on the simiarilty of crop sequences rather than on the longest shared boundary.
This work aims to provide accurate temporal field boundaries for the Contiguous United States through time.
edgesPort.csv: is the processed data derived from https://www.soundtoll.nl/, I cleaned up the data and removed artifacts (such as unclear port names etc...). The exact cleaning process can be found in the R files. The following variables are included:
from_port: port (i) the passage comes FROM
to_port: port (j) the passage is going TO
year: year the relationship is captured
Note: edge weights (frequency of passages between port pairs within each time period) are constructed from the edgelist within the str_net_eda.r file and are not present in edgesPort.csv.
---
db.r: This file constructs the SQL database from the raw files. It is not directly used in the edgelist construction. I am including it for transparency sake.
---
queries.r: This file performs SQL queries to construct the edgelist.
---
str_net_eda.r: This file conducts the network topology used in a forthcoming publication: The Structure of European Maritime Trade: A Network Topology of the Sound Toll Registers, 1670-1857 (sent out for review)
SD spectroscopy folder: the raw h5py files stored in the folder with date when the experiment taken / recorded RID(experiment name) and average mean counts over 20 repeitions in Excel files (SD spectroscopy RID only1 and SD spectroscopy) / SD spectroscopy.py for data analysis.
Shelving rate: raw h5py file in the folder with date when the experiment taken / recorded RID (experiment name) in Excel file (411 Optical pumping) / Shelving rate.py for data analysis
Shelving experiment: raw h5py file in the folder with date when the experiment taken / recorded RID (experiment name) in Excel file for each case of shelving (Shelving data1 for three ion dynamics: each column for each shelving configuration and a group of 4 RIDs corresponds to single time point & Two ion shelving exp for two ion dynamics: first column as time point, the others are RID at that time point) / Three ion shelving experiment data analysis.py and Two ion shelving experiment data analysis.py for data analysis.
Deshelving rate: raw h5py file in the folder with date when the experiment taken / recorded RID (experiment name) in Excel file for each Rabi frequency case (Deshelving RID xx kHz) / Deshelving rate.py for data analysis
Citation to related publication:
arXiv preprint arXiv:2602.10307
Title:
Dataset for "In-Situ Rewiring of Two-Dimensional Ion Lattice Interactions Using Metastable State Shelving"
The stimuli consisted of 200 items, among which 100 critical items included 25 quadruples as in (1a-d), with 50 more English-like items in which the gender of the clitic or strong pronoun reflected the gender of the antecedent (1a, b), and 50 less English-like items in which the genitive pronoun agreed in gender with the head noun of the genitive structure rather than with the antecedent (1c, d). Of these second 50 items, 25 had a female antecedent in the matrix clause as in (1c), and 25 had a male antecedent as in (1d). Similarly, half of the more English-like items (1a, b) involved masculine pronouns le and lui ‘3p.sing.masc’ and half involved feminine pronouns la and elle ‘3p.sing.fem’. Crucially, antecedent-gender-specified pronouns la and elle and antecedent-gender-unspecified pronoun son all allow the retrieval of the matrix subject as the antecedent. The 100 distractor items involved complex interrogative structures and permutations like target items, counterbalanced so that no grouping stood out.
(1a) Quelle décision le concernant est-ce que Paul a dit t que Lydie avait rejetée t
which decision him regarding is-it that Paul has said that Lydie had rejected
sans hésitation ?
without hesitation
‘Which decision regarding him did Paul say that Lydie had rejected without hesitation?’
(1b) Quelle décision à propos de lui est-ce que Paul a dit t que Lydie avait rejetée t
which decision about him is-it that Paul has said that Lydie had rejected
sans hésitation ?
without hesitation
‘Which decision about him did Paul say that Lydie had rejected without hesitation?’
(1c) Quelle décision à son sujet est-ce que Paul a dit t que Lydie avait rejetée t
which decision about him/her is-it that Paul has said that Lydie had rejected
sans hésitation ?
without hesitation
‘Which decision regarding him did Paul say that Lydie had rejected without hesitation?
(1d) Quelle décision à son sujet est-ce que Lydie a dit t que Paul avait rejetée t
which decision about him is-it that Lydie has said that Paul had rejected
sans hésitation ?
without hesitation
‘Which decision about him did Lydie say that Paul had rejected without hesitation?’
E-Prime delivered the stimuli in a rapid serial visual presentation (RSVP) reading task. The stimuli appeared word by word at the center of the screen in 36-point Consolas font, using normal orthographic conventions. They appeared in four blocks presented in random order. Within each block, stimuli were also presented in random order. Participants sat in a chair facing a computer monitor at a distance of approximately four feet. A fixation cross at the center of the screen preceded each item, lasting 700ms. The task was found to be hard but manageable to advanced L2 speakers in stimuli preparation. Due to the time required for E-prime to load each word and for the monitor’s refresh rate, the total presentation time per word was 566ms (300ms presentation, 250ms interstimulus interval, and a 16ms refresh rate between words) accommodating L2 speakers without being unnaturally slow for L1 speakers.
Respondents were trained to read questions like the stimuli and then accept or reject follow-up comprehension statements, which were presented in their entirety for a maximum of 3500ms. These comprehension checks were of several types: Some examined the propensity for an anaphoric interpretation, while others queried other aspects of the sentences. Participants quickly responded to the statements by pressing the left arrow key for ‘Yes/True’ and the right arrow key for ‘No/False.’ There was a training session of six items, which could be repeated before moving on to the experiment. In the training, all items were followed by a comprehension statement; in the task, only two thirds were. However, L1 and L2 speakers alike interpreted the pronoun as referring to the gender-matched noun phrase 70% of the time in critical stimuli. Naturally, a set of questions like our stimuli seems plausible in only a limited set of situations. Thus, respondents were introduced to a context involving two characters who were devoted followers of a television series. One of the characters, however, had missed some episodes and asked the other character questions to catch up.
This dataset was generated as part of a multi‑year research effort examining differences in autism spectrum disorder (ASD) knowledge among caregivers and providers across disciplines and levels of experience . This dataset consists of de‑identified survey data from the Autism Knowledge Survey–Revised (AKS‑R) and is provided in SPSS (.sav) format with variable names, labels, and value labels embedded. The file includes item‑level responses, derived knowledge scores, and basic demographic variables related to participant role and experience. No direct or indirect personal identifiers are included.
Data was compiled from several publicly available sources of scientific and educational data (NSF Awards Search, SciSciNet data lake, Web of Science broad disciplinary, Lightcast skills, O*Net Content Model Reference) with details and links available in the dataset readme. Text from these data sources was processed using NLP methods in a series of Python and R scripts, available with full description in the dataset repository at
https://github.com/monica-marion/Interdisciplinary-Training-Grants-Collection.
The core data set are citations, typically in a basically Scopus dataset format, for 92 reports published between 1987 and 2025. Included are various views of the data set and tables summarizing some of the characteristics of the data set. Also included are a few sheets explaining calculations that do one of two things: convert Return On Investment (ROI) data from the form initially presented in a report to the units used in our analyses; or how Return on Investment figures were calculated from reports that included enough data to enable ROI calculation but within which an ROI figure was not calculated.
Snapp-Childs, W., Hancock, D., Smith, P., Towns, J., Stewart, C. (2025) "Overview of best practices for quantitative analysis of economic and academic benefits of research-enabling facilities." Indiana University. https://hdl.handle.net/2022/34680. https://hdl.handle.net/2022/34680
Title:
PRISMA Data and Analyses Regarding University-based Research-enabling Facilities and ROI
Jupyter Notebook pages that include automation protocols were created for each step in the the multistep continuous-flow synthesis of a D-glucuronic acid building block through a series of optimised and modular transformations, including O-p-methoxyphenyl (PMP) glycosylation, Zemplén deacylation, 4,6-O-benzylidene protection/deprotection, 2-O-benzoylation, and C6 oxidation/methylation. The Jupyter Notebooks serve as a versatile platform for both writing Python code and comprehensively documenting automated chemical procedures, including reaction setups, protocols, and execution logs. To ensure experimental reproducibility, the necessary chemical information is stored externally in JSON files. The Notebooks process this data to dynamically generate stoichiometry tables and apparatus descriptions, while simultaneously controlling the laboratory’s liquid handling systems.
This collection documents open-source Python-based code used to build synthesizers and execute synthetic protocols to produce various chemical compounds.
The data was collected through an online Qualtrics survey. R and R Studio were used to subset the variables of interest that can be used in statistical analysis. No specific software or scripts, however, are required to access the final CSV file.
In addition to participant demographics, the number and type of herbs and spices used, supplements taken, specific diseases had, and number of prescription medications taken where analyzed in this dataset. Only this select information on the forms was loaded into the dataset. SPSS was used to run the dataset statistics.
This paricular collection contains namelist.input, cape.zip, radar.zip, precip.zip, surface.zip, updraft_helicity.zip, vorticity.zip, xsec.zip, and wrfout_d01_2010-05-02_13_00_00.nc.namelist is configuration file of WRF. cape is short for Convective Available Potential Energy, a measure of the instability in an air mass. cape.zip is the visualization of cape and contains 24 png files. radar is Mix of radar minimum and radar maximum visualizations. radar.zip represents the mixed results of putting those two radar types together. radar.zip is the visualization of vorticity and contains 28 png files. precip is short for Precipitation, the sum of the rain, snow and hail in given in liquid equivalent depth. precip.zip is the visualization of precip and contains 4 png files. surface is meteorological parameters on the earth's surface, or in a model on the first level above the ground. surface.zip is the visualization of surface and contains 16 png files. updraft_helicity is the dot product of the vertical velocity and the vertical vorticity. It is presented as a summation over a 3-km depth. updraft_helicity.zip is the visualization of updraft_helicity and contains 16 png files. vorticity is the localized rotation of the air. In model plots it is often the vertical component of vorticity, the rotation of the horizontal winds. vorticity.zip is the visualization of vorticity and contains 16 png files. xsec is is the cross section. xsec.zip is the visualization of xsec and contains 51 png files. wrfout_d01_2010-05-02_13_00_00 is computational result of WRF.
Title:
Vortex II Forecast Data - forecast_20100519150000Z_run001
All data was processed in Microsoft Excel. An exception is the TIRF microscopy data, which led to Microsoft Excel for final computations and building the histogram. Across all data, measurements were taken from either the absorbance or the fluorescence of a fluorescent molecule (CAM-ALEXA488, 496/515 nm) and capsid protein (280 nm).