The consulting firm is founded by four computer science experts from Saarland University to provide sound advice on data analysis. Data Science Consulting focuses on collecting, cleaning and merging data, as well as removing errors, maintaining data, setting up a scalable architecture and defining critical characteristics for analysis.
The FraudBuster approach uses proactive risk prediction at the underwriting stage to identify unprofitable drivers who are likely fraudulent risks. This novel method can help insurers reduce fraud in high-risk markets, such as the automobile insurance market.
A study published in EPJ Data Science found that fiction books, especially those with high initial sales numbers, are more likely to become bestsellers. Non-fiction titles, however, are more likely to retain their bestseller status once achieved.
Portland State University has received a $300,000 grant to bring eight college students from across the U.S. to work on real-world projects using computational modeling and big data to improve Portland's quality of life. The program aims to provide valuable mentoring and research experience to undergraduate students.
GW and UGA researchers develop GlyGen portal to integrate glycan data with gene and protein data, enabling more effective analysis. The project simplifies the process of understanding glycans' roles in diseases like cancer.
Researchers have turned salmon migration patterns into sound using sonification, enabling untrained listeners to interpret large amounts of complex data. The approach has shown promise in helping scientists feel less overwhelmed by interpreting big data, leading them to spend more time exploring the experience.
A team from Lehigh University's Industrial and Systems Engineering department participated in a workshop on big data optimization algorithms, theory, applications and systems. Researchers presented their work on Interior Point Methods, machine learning methodologies, and other topics relevant to big data analytics.
Researchers employed computational approaches to estimate the fitness landscape of gp160, a polyprotein that comprises HIV's spike. The inferred landscape was validated through comparisons with diverse experimental measurements.
Experiments suggest that farsightedness is negatively associated with risk-taking and positively associated with choosing increased future rewards. Analysis of over 90 million tweets reveals a connection between future thinking and decision-making.
A new algorithm, VarQuest, has been developed to identify naturally occurring antibiotics, revealing an order of magnitude more compounds than previous studies. This breakthrough could lead to the discovery of new medicines and help combat rising antibiotic resistance.
A new collection of articles explores the role and risks of bots in influencing public opinion and political elections. Researchers examine approaches to detect and control bots, their impact on recent elections, and potential methods for automation identification of bot activity.
Researchers developed a tool to help users cloak their identity and avoid certain types of inferences drawn about them based on their social media behavior. The 'cloaking device' can significantly reduce the predictive value of online information used for targeting.
A critical review of existing literature reveals that social media big data can be used to monitor and intervene on behalf of people with drug addiction and abuse problems. Researchers have developed an evidence-based framework to inform future social media-based substance use prevention and recovery programs.
The UW Department of Astronomy is joining the ZTF team to develop new methods for identifying celestial objects in the night sky. The ZTF's massive real-time data stream will impact studies of stars, our solar system, and the evolution of our universe.
A North Carolina State University-led study found that six imperatives facilitate collaboration in novel partnerships between government, academia, and industry. These include mutual benefit, trust, common understanding, access and agility, incentives, and mission criticality.
The Data Science Institute developed a novel statistical method to measure predictivity in big data analysis. The approach allows researchers to compare their predictions to a theoretical baseline, enhancing accuracy. The team will help the New York City Department of Transportation assess complex social problems using big data sets.
The traditional 24 Solar Terms are being upgraded with new, geographically correlated models using big meteorological data. The updated system reveals inconsistent timing and spatial inhomogeneity between the old and new systems.
Researchers found that big data analytics can amplify existing policing practices, leading to more marginalization and distrust. The influx of personal data enables law enforcement to surveil communities more easily, but also raises concerns about the use of objective crime data and predictive algorithms.
Researchers have developed a new algorithm that allows AI to collect error reports and correct them immediately, without affecting existing skills. This enables robots to learn from their mistakes and spread new knowledge amongst themselves.
A groundbreaking study applies big data analysis to mineralogy, predicting the existence of 1,500 missing minerals and new deposits. The technique enables scientists to represent data from multiple variables on thousands of minerals in a single graph, revealing patterns of occurrence and distribution.
A new special issue of Big Data highlights the risks associated with big data, including discrimination, lack of diversity, and bias. The issue discusses strategies for making decision-making 'discrimination-aware' and emphasizes the importance of considering ethical issues in model development.
Researchers from Penn State's IST have developed a method to identify bow echoes in radar images, a phenomenon associated with fierce winds. The algorithm can automatically detect bow echoes as they begin to form, providing instant notifications for severe weather alerts.
Researchers at Tsinghua University outline recent advances on nonparametric Bayesian methods, regularized Bayesian inference, scalable algorithms, and system implementation to tackle the challenges of Big Data. They also discuss connections with deep learning and highlight the need for human expertise in devising appropriate features a...
Researchers found that simple 'single-show models' can have high predictive accuracy in predicting presidential election outcomes based on television viewership data. This approach may offer a more accurate predictive tool compared to poll-data-driven models.
Industry experts examine the advantages of shifting data analytics to the cloud, including cost savings and improved management of analytics workloads. The panel discussion highlights the potential for companies to gain by moving their data analytics activities to the cloud, including more efficient use of IT resources.
Researchers developed an adaptive control approach based on online learning to correct dynamics errors in real-time, improving robustness of motion systems. The DOOMED algorithm updates a correction model until correct acceleration is achieved, minimizing error between desired and actual accelerations.
A new KIT Motion-Language Dataset has been created to support the development of robot activities based on natural language input. The dataset, which includes over 4,000 motions and 6,200 annotations in natural language, aims to unify and standardize research linking human motion and natural language.
Xia emphasizes that big data is about more than just numbers and requires new mathematical methods to analyze. Researchers in mathematics, signal processing, and computer science must develop these new tools to unlock the full potential of big data.
Data Civilizer aggregates scattered data from various files, creating unified datasets for analysis. The system identifies commonalities between columns and traverses a map to find related data, enabling users to compose queries and save results.
Computing science researchers at the University of Alberta have developed a technique to automate geotagging for news articles and other online documents. The model integrates two competing hypotheses: inheritance and near-location, achieving high accuracy in matching named entities to geographical locations.
A Princeton-led team has created a new measure called the influence score, which can effectively differentiate between noisy and predictive variables in big data. This approach significantly improves prediction rates in various fields, including breast cancer diagnosis, terrorism, and financial markets.
Researchers identified six core emotional storylines: rags to riches, richness to rags, man in a hole, icarus, Cinderella, and Oedipus. These findings may help create compelling stories and teach common sense to AI systems.
The NIH-led initiative reviews the use of big data in infectious disease surveillance, combining traditional and non-traditional data sources to provide more accurate and timely information. This approach aims to forecast outbreak sizes and trajectories, enabling better responses to emerging threats.
New research using big data analysis has found that people's collective behaviour is more predictable than thought, with strong periodic patterns influenced by the weather and seasons. Historical news and social media data revealed cycles in leisure and work activities, diet, diseases, and mental health.
Researchers developed a way to integrate multiple big data sets from biology to understand cellular processes, discovering new regularities and biological consistencies. The study found pause sites dictate protein structure and folding, providing insights into cancer biology.
Interdisciplinary research highlights changing scientific landscape, where large data sets and computational methods encourage an iterative approach. The authors note that despite new technology, the reinvigorated approaches are rooted in centuries-old debates over iterative and hypothesis-driven science.
A novel tensor mining tool enables automated modeling in big data applications, facilitating the analysis of complex multiaspect data. This innovation addresses the challenge of extracting knowledge from massive amounts of data represented as tensors.
A new data-cleaning tool called ActiveClean analyzes a user's prediction model to identify mistakes and update the model as it works. By minimizing human error, ActiveClean improves model accuracy and reduces statistical biases, making it an essential tool for building better prediction models.
Researchers are uniting to tackle the complex challenge of understanding brain function through large-scale computational modeling. This approach aims to improve our knowledge of brain function by creating realistic models based on biological data.
A Bristol student, Paul Harris, has achieved a world record in 5G wireless spectrum efficiency using Massive MIMO. He set a new record of 145.6 bit/s/Hz with his research team, demonstrating the potential for this technology to deliver ultra-fast data rates to high densities of smartphones and tablets.
Researchers have identified comprehensibility as a key goal in model development, considering stakeholders' understanding of the modeling process. The article provides a holistic framework for comprehensibility in data science projects, prioritizing human needs and understanding.
Researchers developed a novel visualization tool to explore dynamic Bitcoin transactional data, revealing behavioral patterns and potential money laundering activities. The top-down approach enables drill-down analysis of individual transactions.
Insilico Medicine scientists will present advances in deep learning for biomarker development and drug discovery at the ISFA-Columbia University Actuarial Science Workshop. The workshop aims to integrate deep learning with actuarial science to assess risk in finance, insurance, and other industries.
The Global Names project enables the rapid indexing of content using scientific names, improving the discovery of small data. The study found that name-matching was improved to almost 85% with simplified or canonical versions of names.
A novel type 2 diabetes risk model has been developed to better understand disease progression, revealing people on atypical trajectories face significantly increased or decreased risks of developing T2D. The new model has overcome challenges associated with estimating T2D onset based on comorbid conditions.
The Instituto de Astrofísica de Canarias will participate in the SUNDIAL network, training young researchers in astronomy and computer science to understand galaxy formation and evolution. The network aims to detect ultradiffuse galaxies and apply research to society in medical imaging and remote sensing.
A new doctoral training program at the University of Missouri aims to create a unique type of data scientist by combining expertise from life sciences, medicine, and computing. The six-student program will focus on massive and complex data analytics for one health, using Big Data practices to improve medical discoveries.
A new simulator is being developed to help individuals and families select the most suitable health insurance plans based on realistic cost estimates and potential healthcare outcomes. The simulator combines buyer-specific information with bespoke databases, producing transparent output that enables informed decision-making.
Researchers at University of Houston are developing a new framework to protect big data processing and minimize risks of data breaches on the cloud. The project aims to address security concerns in large-scale data analytics, with potential commercial value in the growing $125 billion market.
A Penn State psychologist argues that big data can enhance our understanding of human development by aggregating empirical work from multiple investigators. This approach could also enable personalized medicine and wearable data-collection with more accurate results.
University of Illinois researchers have achieved record-breaking speeds for fiber-optic data transmission, reaching 57 Gbps at room temperature and 50 Gbps at higher temperatures. This technology could enable faster data transfer and use of large data streams in applications such as virtual reality.
The symposium highlights opportunities to utilize federal big data initiatives in dental research, such as BD2K, to advance research and practice. Panelists will discuss the potential of linked data for dental providers.
RevEx performs faceted searches and analyzes text and data across multiple domains to reveal important findings. Its applications range from investigating medical services to visualizing humanitarian data on a country-by-country basis.
Researchers developed Eyebrowse, a system allowing users to share self-selected aspects of their online activity with friends and the public. Users can add whitelisted sites, track friend visits, and view community browsing history, providing insights for academics and companies targeting consumers.
A new study proposes using data mining tools to identify patterns in QS data that can inform users' decisions on diet and exercise. This approach has the potential to reveal ways to improve personal well-being without compromising data privacy.
A national patient-powered registry, MyLymeData, has enrolled over 3,000 patients to accelerate research for chronic Lyme disease. Big data tools can help identify treatment responses and subgroup analysis, leading to a better understanding of the disease.
A new study identified distinct patient diagnoses and emergency department usage patterns linked to high risk of ED readmission. The researchers used electronic healthcare records data from over one million patients to build predictive models for risk of 72-hour ED readmission.
Big data analytics is revolutionizing healthcare by developing predictive models for disease onset, such as type 2 diabetes, and providing personalized insights through wearable devices. The new approaches aim to improve diagnosis accuracy and patient outcomes.
Recent epigenetics research highlights molecular mechanisms influencing gene expression through socioenvironmental factors. Big data raises concerns over immortalized participant data, privacy, and anonymity, prompting recommendations for strengthening ethical consent practices.
Scientists used a database of 418 plant species to analyze patterns in growth, reproduction, and survival. They found that two characteristics - growth rate and reproduction strategy - can predict population growth and response to disturbances.