<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>data science &#8211; Binghamton University Research News</title>
	<atom:link href="https://discovere.binghamton.edu/tag/data-science/feed" rel="self" type="application/rss+xml" />
	<link>https://discovere.binghamton.edu</link>
	<description>Insights and Innovations From Binghamton University</description>
	<lastBuildDate>Thu, 08 Dec 2022 14:43:09 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	
	<item>
		<title>Student analyzes tweets to gauge political sentiment</title>
		<link>https://discovere.binghamton.edu/student-spotlights/foreman-8325.html</link>
		
		<dc:creator><![CDATA[Tasfia Rubayat]]></dc:creator>
		<pubDate>Mon, 09 Jan 2023 13:00:11 +0000</pubDate>
				<category><![CDATA[Students]]></category>
		<category><![CDATA[data science]]></category>
		<category><![CDATA[political science]]></category>
		<category><![CDATA[politics]]></category>
		<category><![CDATA[Russia]]></category>
		<category><![CDATA[Ukraine]]></category>
		<guid isPermaLink="false">https://discovere.binghamton.edu/?p=8325</guid>

					<description><![CDATA[Binghamton undergraduate Lisa Foreman incorporates her passions for political science and Russian studies into her research. ]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" class="alignleft size-full wp-image-8337" src="https://discovere.binghamton.edu/wp-content/uploads/2022/12/foreman_02.jpg" alt="" width="132" height="133" srcset="https://discovere.binghamton.edu/wp-content/uploads/2022/12/foreman_02.jpg 132w, https://discovere.binghamton.edu/wp-content/uploads/2022/12/foreman_02-120x120.jpg 120w" sizes="(max-width: 132px) 100vw, 132px" />Using tweets from the accounts of influential political figures, a Binghamton undergrad analyzes sentiment in order to read between the lines.</p>
<p>Lisa Foreman has been an active participant in multiple research programs on campus since her freshman year. Now, in her third year, she has effectively incorporated her passions for political science and Russian studies into her research.</p>
<p>Foreman was able to pave her own path with her studies, thanks in large part to the Source Project, a first-year honors research program in which students can choose between various streams of study, ranging from education to sustainability in art, to complete research in the humanities.</p>
<p>“When I saw other people doing research, I was like ‘Oh man! I need to start researching!’ Foreman says. “By the time I was accepted to Binghamton and to the Source Project, I was already so excited because this was what I was waiting for, now it was my turn to join the scholarly conversation.”</p>
<p>Foreman, who grew up on Staten Island, developed her core research interests as a first-year student. She also credits the program with allowing her to continue research in Russian studies.</p>
<p>“Before I had even accepted my enrollment to this school, I was already attending info sessions and learning about what it would be like to be in the specific track that I was interested in,” she says. “I felt like I was appreciated immediately, and I felt like I was really understood as an applicant.”</p>
<p>In addition to her research projects, Foreman is involved with TEDxBinghamton University. As the director of content, she plays a significant role in organizing the annual conference in the spring, and encouraging student speakers, alumni speakers and professionals who want to share their message on the TEDx stage.</p>
<p>“We tap into every corner of the University possible, while also bringing something new to them,” Foreman says.</p>
<p>During the summer of 2022, Foreman took part in the Summer Scholars and Arts Program, conducting research in sentiment analysis in relation to the Russian-Ukrainian conflict. She was referred to the program by a former professor, Mikhail Filippov, who now serves as her advisor.</p>
<p>“We have already started to work on her 80-page honors thesis and I have no doubts that she will complete it successfully,” says Filippov, a professor of political science. “Overall, it is a very rewarding experience for me, largely because Lisa acts as a highly motivated, excellent student with self-discipline.”</p>
<p>Foreman uses a research plan, inspired by her work in the Source Project, to analyze tweets selected from the Twitter accounts of influential political leaders such as President Joe Biden and U.S senators.</p>
<p>Foreman identifies key words and phrases that correlate with sentiment. A single tweet is analyzed for its polarity, as in whether it is positive or negative, as well as the degree at which this is expressed, which is considered as its subjectivity. For example, a tweet with positive polarity and high subjectivity would have a positive sentiment.</p>
<p>As she outlines the background, methods and question of inquiry in her fundamental work, Foreman is able to improve the guidelines and accuracy of her project.</p>
<p>“By unpacking certain cross sections of social media, information and conflict, my research can be used to expand awareness regarding political social media,” she says. “I amalgamated so much knowledge and understanding of this process, but things can definitely be narrowed.”</p>
<p>Foreman says she has had a unique experience studying Russian literature during a time of war in the region. Through her classes, Foreman experienced the Russian-Ukrainian conflict from a humane and scholarly perspective.</p>
<p>“I was in a Russian language course at that point, and one of my classmates was Ukrainian,” she says. “It was spoken about a lot within the Russian department, and even outside of that. It was interesting to not only discuss it in a casual sense, but to really be exposed to people who were directly affected by it, while also considering it as an academic.”</p>
<p>The application of computer programming in her project introduced a new aspect that allowed for greater creativity, Foreman says. “With an essay you write until you feel like your argument is complete, but that&#8217;s kind of ambiguous,” she says. “Coding felt more artistic and creative, more so than any other math or science class I had taken before, as well as how much you can do with it.”</p>
<p>Foreman intends to explore the role of social media in conflict and how computer-aided text analysis can be used to understand it. She hopes that her project will contribute to the growing discussion of how social networking platforms present global affairs.</p>
<p>“Social data science is an emerging discipline that my project fits into,” Foreman says. “I’m just trying to keep up with the curve and contribute something to this discussion while tying it into another fairly new discipline.”</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Machine learning research may aid industry</title>
		<link>https://discovere.binghamton.edu/student-spotlights/banihani-7564.html</link>
		
		<dc:creator><![CDATA[Jacob T. Kerr]]></dc:creator>
		<pubDate>Tue, 03 Dec 2019 14:00:33 +0000</pubDate>
				<category><![CDATA[Students]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[data]]></category>
		<category><![CDATA[data science]]></category>
		<category><![CDATA[machine learning]]></category>
		<guid isPermaLink="false">https://discovere.binghamton.edu/?p=7564</guid>

					<description><![CDATA[A graduate student has created an "oracle" that can make accurate predictions related to spam emails, bank fraud, workers quitting their jobs and more. ]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" class="alignleft size-full wp-image-7590" src="https://discovere.binghamton.edu/wp-content/uploads/2019/10/bani_hani_02.jpg" alt="" width="132" height="133" />Spam emails, bank fraud, diabetes, workers quitting their jobs. What do these topics have in common? The answer can be found in machine learning research at Binghamton University.</p>
<p>Dana Bani-Hani, a doctoral student studying industrial and systems engineering, has spent the past few years teaching machines how to read data sets in any industry. The system she coded, called a Recursive General Regression Neural Network Oracle (R-GRNN Oracle), takes data inputs and creates prediction outputs.</p>
<p>Classification models are not new in data science and analytics, but what Bani-Hani created goes beyond the basics. A typical system uses algorithms, called classifiers, that run through a data set of many different variables to create a prediction. Oracles are created to run multiple sets of these classifiers to see which algorithm creates the most accurate prediction.</p>
<p>For example, a classifier can look at a myriad of emails and factor in certain word usage, word count and several other variables to determine if the email is spam. An oracle looks at the different classifier outputs and determines which most accurately predicted the spam emails.</p>
<p>What sets the R-GRNN Oracle apart from other oracles is its capability to take classifier outputs and rank them based on their accuracy. Based on the ranking, classifiers are given weights and are combined to produce a prediction superior to any one classifier on its own.</p>
<p>Think of this process like an orchestra. Each instrument has its own strengths, just like different classifiers, so it is useful to include them all. The conductor, like the R-GRNN Oracle, directs the different instruments to play loudly or more softly based on how the instrument makes the final symphony sound.</p>
<p>At this point, the system would be called a General Regression Neural Network (GRNN), which has been created before at Binghamton University. The real crux of Bani-Hani’s work lies in the first letter, R, standing for Recursion.</p>
<p>The R-GRNN Oracle takes the original GRNN output, and uses that entire system as an input for another GRNN prediction. This is combined with the most successful of the original classifiers.</p>
<p>So, back to the orchestra: The original symphony is recorded, and then played back again later. This time, along with the recording, a few instruments play again to further fine-tune the important sounds of the orchestra.</p>
<p>“Because of the way [the GRNN] works, I was able to create the recursive model,” Bani-Hani says. “The concept of recursion is not widely used in machine learning, so I decided to put an oracle inside of an oracle.”</p>
<p>Mohammad Khasawneh, professor and department chair in systems science and industrial engineering, supervised Bani-Hani’s research. He says systems like the GRNN and R-GRNN are underutilized and are vital in serious life events.</p>
<p>“The traditional GRNN Oracle has received limited attention in the literature as only very few researchers have published work on the algorithm,” Khasawneh says. “But many real-life problems that apply machine learning models to automate classifying unknown observations require accurate predictions. Tasks such as diagnosing diseases entail precision to avoid serious issues that could potentially lead to problems such as lawsuits or even deaths.”</p>
<p>Bani-Hani says the R-GRNN Oracle produces more accurate predictions than any single classifier alone, as well as one GRNN on its own. The R-GRNN Oracle took in thousands of email samples, programmed to factor 57 variables, and then produced a spam prediction superior to all other classifiers tested.</p>
<p>Bani-Hani also used the R-GRNN to predict credit card application fraud, diabetes diagnosis and whether a worker will quit based on past workplace experiences. In each case, the R-GRNN came out as the most accurate predictor.</p>
<p>She plans to focus her model on specific fields, such as business or finance, as well as package both the GRNN Oracle and the R-GRNN Oracle so companies do not have to create the entire code from scratch.</p>
<p>Bani-Hani’s journey to machine learning research started nearly 6,000 miles away from Binghamton in Jordan. After completing her bachelor’s degree in architectural engineering, she heard about Binghamton University through Watson School faculty and academic leaders, and from her father’s supportive suggestions. She initially pursued a master’s degree in industrial engineering, but she soon found a new passion: data mining and machine learning.</p>
<p>&#8220;Getting a PhD has been a dream of mine for the last 15 years,&#8221; Bani-Hani says. &#8220;I mainly attribute this to having a family with advanced degrees. I am thankful to my professors here at Binghamton University for introducing me to the topics that make up my research.&#8221;</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Innovation Day delves into Big Data</title>
		<link>https://discovere.binghamton.edu/news/bigdata-5753.html</link>
		
		<dc:creator><![CDATA[rcoker]]></dc:creator>
		<pubDate>Mon, 28 Apr 2014 15:01:44 +0000</pubDate>
				<category><![CDATA[News]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[data science]]></category>
		<category><![CDATA[innovation]]></category>
		<guid isPermaLink="false">http://discovere.binghamton.edu/?p=5753</guid>

					<description><![CDATA[Innovation Day 2014 at Binghamton University focused on Big Data and how to make sense of it. ]]></description>
										<content:encoded><![CDATA[<p><span style="line-height: 1.5em;"><a href="http://discovere.binghamton.edu/wp-content/uploads/2014/04/innovation_day_2014.jpg"><img fetchpriority="high" decoding="async" class="alignleft size-medium wp-image-5756" src="http://discovere.binghamton.edu/wp-content/uploads/2014/04/innovation_day_2014-300x173.jpg" alt="innovation_day_2014" width="300" height="173" srcset="https://discovere.binghamton.edu/wp-content/uploads/2014/04/innovation_day_2014-300x173.jpg 300w, https://discovere.binghamton.edu/wp-content/uploads/2014/04/innovation_day_2014.jpg 440w" sizes="(max-width: 300px) 100vw, 300px" /></a>Every two days, the world generates as much data as there was in existence until 10 years ago. And while this data has enormous potential, it sometimes feels like we’re drowning in it.</span></p>
<p><span style="line-height: 1.5em;">That was just one of the observations shared during the third-annual Innovation Day at Binghamton University. The April 24 event, which drew about 125 participants from industry and academia, focused on Big Data and how it relates to fields ranging from healthcare to finance. The program, organized by the Office of Entrepreneurship and Innovation Partnerships in conjunction with Binghamton’s Center of Excellence, featured plenary lectures and panels as well as tours of the Innovative Technologies Complex and a research poster session.</span></p>
<p><span style="line-height: 1.5em;">Hao Wang, vice president of information services and chief information officer of the Research Foundation for SUNY, kicked off the day’s lectures by noting that — in some ways — Big Data is nothing new.</span></p>
<p><span style="line-height: 1.5em;">Wang, who used to teach public health, said Big Data can be traced back to London’s cholera epidemic in the 1850s. Dr. John Snow did visualization work by hand, plotting deaths over time on a map of London. He lacked today’s advanced instrumentation and computer modeling, but Snow was able to reveal the source of the cholera outbreak. “Don’t be afraid of Big Data,” Wang said. “… The application is profound. You don’t need microbiology to understand this. Big Data can show you the way.”</span></p>
<p><span style="line-height: 1.5em;">Today, Wang envisions Big Data being used to protect the financial system, to enable drug discovery and to improve public safety, among other applications.</span></p>
<p><span style="line-height: 1.5em;">Katharine Frase, vice president and chief technology officer, IBM Public Sector, spoke about extracting real value from Big Data. Even if it doesn’t feel like it, Big Data affects all of us every day, she said.</span></p>
<p>Frase discussed a recent IBM survey of CEOs that found a third make decisions based on untrustworthy data and that half lack information they need. At the same time, 60 percent said they have too much data. “This,” she said, “is in some ways the <i>real</i> Big Data problem.”</p>
<p>Big Data is often defined by the four Vs, she noted: volume, variety, velocity and veracity.</p>
<p><span style="line-height: 1.5em;">Although reporting — a major source of Big Data — is important, Frase noted that it’s vital to be “more right, more often.” Reporting can be like driving while looking in the rear view mirror, she said. It’s more interesting, not to mention safer, to look through windshield and see what’s coming.</span></p>
<p><span style="line-height: 1.5em;">While experts often discuss intangible benefits of Big Data, it has important implications in the physical world, Frase said. For instance, ConocoPhillips reported $1 billion in savings from accurately predicting ice floes, which resulted in better protection of its oil rigs.</span></p>
<p><span style="line-height: 1.5em;">She offered a few additional examples of businesses getting value from Big Data projects:</span></p>
<ul>
<li>Sprint saw a 90 percent increased transaction capacity.</li>
<li>Qualcomm reported 60 times faster query performance.</li>
<li>MoneyGram saw a 72 percent reduction in fraudulent claims.</li>
</ul>
<p>Frase said Big Data will produce traditional return on investment, or ROI, in some cases. In other instances, the scenario is more complex. For instance, resolving a problem related to traffic may also alleviate pollution and reduce accidents.</p>
<p><span style="line-height: 1.5em;">She suggested that data scientists will have the “sexiest job of the 21st century.” These jobs, she said, will require people who have the domain expertise to ask useful questions as well as the math and computer science expertise to put Big Data to work.</span></p>
<p><span style="line-height: 1.5em;">“Technology is really good at finding answers,” Frase said, “but the humans ought to be the ones asking the questions.”</span></p>
<p><span style="line-height: 1.5em;">In the afternoon, Innovation Day featured a self-proclaimed Big Data naysayer: Scott L. Zeger, professor and vice provost at Johns Hopkins University. “Where there’s too much data, there can be more confusion than clarity,” he said at the outset of his talk on Big Data in health.</span></p>
<p><span style="line-height: 1.5em;">He presented a reworked version of a song called “Help the Poor” by B.B. King and Eric Clapton. With a little assistance from YouTube, Zeger soon had everyone in the Center of Excellence’s new symposium hall singing along:</span></p>
<p><i>Help the poor<br />
</i><i>Won’t you help poor me<br />
</i><i>Got Big Data blues, baby<br />
</i><i>Need your help desperately<br />
</i><i>Big data’s the rage but it’s hard to use well<br />
</i><i>Got to get onboard before that trend does quell<br />
</i><i>Help the poor<br />
</i><i>Oh baby, won’t you help poor me</i></p>
<p><span style="line-height: 1.5em;">Many countries are healthier than we are, Zeger said. Yet the United States spends a trillion dollars more per year on health than Norway, the country that spends the second-largest amount. That trillion dollars? That’s the Big Datum, the one that matters the most, he says.</span></p>
<p><span style="line-height: 1.5em;">Zeger outlined some problems of the American system, chief among them a healthcare system that too often “pays for doing” rather than paying for results. We do a lot more than is done elsewhere, especially for patients with chronic conditions, he said, even though it often doesn’t make people healthier.</span></p>
<p><span style="line-height: 1.5em;">“Big Data,” he said, “can help us cut into that trillion dollars.”</span></p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
