<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>big data &#8211; Binghamton University Research News</title>
	<atom:link href="https://discovere.binghamton.edu/tag/big-data/feed" rel="self" type="application/rss+xml" />
	<link>https://discovere.binghamton.edu</link>
	<description>Insights and Innovations From Binghamton University</description>
	<lastBuildDate>Fri, 06 Dec 2019 15:33:22 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	
	<item>
		<title>Machine learning research may aid industry</title>
		<link>https://discovere.binghamton.edu/student-spotlights/banihani-7564.html</link>
		
		<dc:creator><![CDATA[Jacob T. Kerr]]></dc:creator>
		<pubDate>Tue, 03 Dec 2019 14:00:33 +0000</pubDate>
				<category><![CDATA[Students]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[data]]></category>
		<category><![CDATA[data science]]></category>
		<category><![CDATA[machine learning]]></category>
		<guid isPermaLink="false">https://discovere.binghamton.edu/?p=7564</guid>

					<description><![CDATA[A graduate student has created an "oracle" that can make accurate predictions related to spam emails, bank fraud, workers quitting their jobs and more. ]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" class="alignleft size-full wp-image-7590" src="https://discovere.binghamton.edu/wp-content/uploads/2019/10/bani_hani_02.jpg" alt="" width="132" height="133" />Spam emails, bank fraud, diabetes, workers quitting their jobs. What do these topics have in common? The answer can be found in machine learning research at Binghamton University.</p>
<p>Dana Bani-Hani, a doctoral student studying industrial and systems engineering, has spent the past few years teaching machines how to read data sets in any industry. The system she coded, called a Recursive General Regression Neural Network Oracle (R-GRNN Oracle), takes data inputs and creates prediction outputs.</p>
<p>Classification models are not new in data science and analytics, but what Bani-Hani created goes beyond the basics. A typical system uses algorithms, called classifiers, that run through a data set of many different variables to create a prediction. Oracles are created to run multiple sets of these classifiers to see which algorithm creates the most accurate prediction.</p>
<p>For example, a classifier can look at a myriad of emails and factor in certain word usage, word count and several other variables to determine if the email is spam. An oracle looks at the different classifier outputs and determines which most accurately predicted the spam emails.</p>
<p>What sets the R-GRNN Oracle apart from other oracles is its capability to take classifier outputs and rank them based on their accuracy. Based on the ranking, classifiers are given weights and are combined to produce a prediction superior to any one classifier on its own.</p>
<p>Think of this process like an orchestra. Each instrument has its own strengths, just like different classifiers, so it is useful to include them all. The conductor, like the R-GRNN Oracle, directs the different instruments to play loudly or more softly based on how the instrument makes the final symphony sound.</p>
<p>At this point, the system would be called a General Regression Neural Network (GRNN), which has been created before at Binghamton University. The real crux of Bani-Hani’s work lies in the first letter, R, standing for Recursion.</p>
<p>The R-GRNN Oracle takes the original GRNN output, and uses that entire system as an input for another GRNN prediction. This is combined with the most successful of the original classifiers.</p>
<p>So, back to the orchestra: The original symphony is recorded, and then played back again later. This time, along with the recording, a few instruments play again to further fine-tune the important sounds of the orchestra.</p>
<p>“Because of the way [the GRNN] works, I was able to create the recursive model,” Bani-Hani says. “The concept of recursion is not widely used in machine learning, so I decided to put an oracle inside of an oracle.”</p>
<p>Mohammad Khasawneh, professor and department chair in systems science and industrial engineering, supervised Bani-Hani’s research. He says systems like the GRNN and R-GRNN are underutilized and are vital in serious life events.</p>
<p>“The traditional GRNN Oracle has received limited attention in the literature as only very few researchers have published work on the algorithm,” Khasawneh says. “But many real-life problems that apply machine learning models to automate classifying unknown observations require accurate predictions. Tasks such as diagnosing diseases entail precision to avoid serious issues that could potentially lead to problems such as lawsuits or even deaths.”</p>
<p>Bani-Hani says the R-GRNN Oracle produces more accurate predictions than any single classifier alone, as well as one GRNN on its own. The R-GRNN Oracle took in thousands of email samples, programmed to factor 57 variables, and then produced a spam prediction superior to all other classifiers tested.</p>
<p>Bani-Hani also used the R-GRNN to predict credit card application fraud, diabetes diagnosis and whether a worker will quit based on past workplace experiences. In each case, the R-GRNN came out as the most accurate predictor.</p>
<p>She plans to focus her model on specific fields, such as business or finance, as well as package both the GRNN Oracle and the R-GRNN Oracle so companies do not have to create the entire code from scratch.</p>
<p>Bani-Hani’s journey to machine learning research started nearly 6,000 miles away from Binghamton in Jordan. After completing her bachelor’s degree in architectural engineering, she heard about Binghamton University through Watson School faculty and academic leaders, and from her father’s supportive suggestions. She initially pursued a master’s degree in industrial engineering, but she soon found a new passion: data mining and machine learning.</p>
<p>&#8220;Getting a PhD has been a dream of mine for the last 15 years,&#8221; Bani-Hani says. &#8220;I mainly attribute this to having a family with advanced degrees. I am thankful to my professors here at Binghamton University for introducing me to the topics that make up my research.&#8221;</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Innovation Day delves into Big Data</title>
		<link>https://discovere.binghamton.edu/news/bigdata-5753.html</link>
		
		<dc:creator><![CDATA[rcoker]]></dc:creator>
		<pubDate>Mon, 28 Apr 2014 15:01:44 +0000</pubDate>
				<category><![CDATA[News]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[data science]]></category>
		<category><![CDATA[innovation]]></category>
		<guid isPermaLink="false">http://discovere.binghamton.edu/?p=5753</guid>

					<description><![CDATA[Innovation Day 2014 at Binghamton University focused on Big Data and how to make sense of it. ]]></description>
										<content:encoded><![CDATA[<p><span style="line-height: 1.5em;"><a href="http://discovere.binghamton.edu/wp-content/uploads/2014/04/innovation_day_2014.jpg"><img fetchpriority="high" decoding="async" class="alignleft size-medium wp-image-5756" src="http://discovere.binghamton.edu/wp-content/uploads/2014/04/innovation_day_2014-300x173.jpg" alt="innovation_day_2014" width="300" height="173" srcset="https://discovere.binghamton.edu/wp-content/uploads/2014/04/innovation_day_2014-300x173.jpg 300w, https://discovere.binghamton.edu/wp-content/uploads/2014/04/innovation_day_2014.jpg 440w" sizes="(max-width: 300px) 100vw, 300px" /></a>Every two days, the world generates as much data as there was in existence until 10 years ago. And while this data has enormous potential, it sometimes feels like we’re drowning in it.</span></p>
<p><span style="line-height: 1.5em;">That was just one of the observations shared during the third-annual Innovation Day at Binghamton University. The April 24 event, which drew about 125 participants from industry and academia, focused on Big Data and how it relates to fields ranging from healthcare to finance. The program, organized by the Office of Entrepreneurship and Innovation Partnerships in conjunction with Binghamton’s Center of Excellence, featured plenary lectures and panels as well as tours of the Innovative Technologies Complex and a research poster session.</span></p>
<p><span style="line-height: 1.5em;">Hao Wang, vice president of information services and chief information officer of the Research Foundation for SUNY, kicked off the day’s lectures by noting that — in some ways — Big Data is nothing new.</span></p>
<p><span style="line-height: 1.5em;">Wang, who used to teach public health, said Big Data can be traced back to London’s cholera epidemic in the 1850s. Dr. John Snow did visualization work by hand, plotting deaths over time on a map of London. He lacked today’s advanced instrumentation and computer modeling, but Snow was able to reveal the source of the cholera outbreak. “Don’t be afraid of Big Data,” Wang said. “… The application is profound. You don’t need microbiology to understand this. Big Data can show you the way.”</span></p>
<p><span style="line-height: 1.5em;">Today, Wang envisions Big Data being used to protect the financial system, to enable drug discovery and to improve public safety, among other applications.</span></p>
<p><span style="line-height: 1.5em;">Katharine Frase, vice president and chief technology officer, IBM Public Sector, spoke about extracting real value from Big Data. Even if it doesn’t feel like it, Big Data affects all of us every day, she said.</span></p>
<p>Frase discussed a recent IBM survey of CEOs that found a third make decisions based on untrustworthy data and that half lack information they need. At the same time, 60 percent said they have too much data. “This,” she said, “is in some ways the <i>real</i> Big Data problem.”</p>
<p>Big Data is often defined by the four Vs, she noted: volume, variety, velocity and veracity.</p>
<p><span style="line-height: 1.5em;">Although reporting — a major source of Big Data — is important, Frase noted that it’s vital to be “more right, more often.” Reporting can be like driving while looking in the rear view mirror, she said. It’s more interesting, not to mention safer, to look through windshield and see what’s coming.</span></p>
<p><span style="line-height: 1.5em;">While experts often discuss intangible benefits of Big Data, it has important implications in the physical world, Frase said. For instance, ConocoPhillips reported $1 billion in savings from accurately predicting ice floes, which resulted in better protection of its oil rigs.</span></p>
<p><span style="line-height: 1.5em;">She offered a few additional examples of businesses getting value from Big Data projects:</span></p>
<ul>
<li>Sprint saw a 90 percent increased transaction capacity.</li>
<li>Qualcomm reported 60 times faster query performance.</li>
<li>MoneyGram saw a 72 percent reduction in fraudulent claims.</li>
</ul>
<p>Frase said Big Data will produce traditional return on investment, or ROI, in some cases. In other instances, the scenario is more complex. For instance, resolving a problem related to traffic may also alleviate pollution and reduce accidents.</p>
<p><span style="line-height: 1.5em;">She suggested that data scientists will have the “sexiest job of the 21st century.” These jobs, she said, will require people who have the domain expertise to ask useful questions as well as the math and computer science expertise to put Big Data to work.</span></p>
<p><span style="line-height: 1.5em;">“Technology is really good at finding answers,” Frase said, “but the humans ought to be the ones asking the questions.”</span></p>
<p><span style="line-height: 1.5em;">In the afternoon, Innovation Day featured a self-proclaimed Big Data naysayer: Scott L. Zeger, professor and vice provost at Johns Hopkins University. “Where there’s too much data, there can be more confusion than clarity,” he said at the outset of his talk on Big Data in health.</span></p>
<p><span style="line-height: 1.5em;">He presented a reworked version of a song called “Help the Poor” by B.B. King and Eric Clapton. With a little assistance from YouTube, Zeger soon had everyone in the Center of Excellence’s new symposium hall singing along:</span></p>
<p><i>Help the poor<br />
</i><i>Won’t you help poor me<br />
</i><i>Got Big Data blues, baby<br />
</i><i>Need your help desperately<br />
</i><i>Big data’s the rage but it’s hard to use well<br />
</i><i>Got to get onboard before that trend does quell<br />
</i><i>Help the poor<br />
</i><i>Oh baby, won’t you help poor me</i></p>
<p><span style="line-height: 1.5em;">Many countries are healthier than we are, Zeger said. Yet the United States spends a trillion dollars more per year on health than Norway, the country that spends the second-largest amount. That trillion dollars? That’s the Big Datum, the one that matters the most, he says.</span></p>
<p><span style="line-height: 1.5em;">Zeger outlined some problems of the American system, chief among them a healthcare system that too often “pays for doing” rather than paying for results. We do a lot more than is done elsewhere, especially for patients with chronic conditions, he said, even though it often doesn’t make people healthier.</span></p>
<p><span style="line-height: 1.5em;">“Big Data,” he said, “can help us cut into that trillion dollars.”</span></p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
