<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>data &#8211; Binghamton University Research News</title>
	<atom:link href="https://discovere.binghamton.edu/tag/data/feed" rel="self" type="application/rss+xml" />
	<link>https://discovere.binghamton.edu</link>
	<description>Insights and Innovations From Binghamton University</description>
	<lastBuildDate>Fri, 06 Dec 2019 15:33:22 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	
	<item>
		<title>Machine learning research may aid industry</title>
		<link>https://discovere.binghamton.edu/student-spotlights/banihani-7564.html</link>
		
		<dc:creator><![CDATA[Jacob T. Kerr]]></dc:creator>
		<pubDate>Tue, 03 Dec 2019 14:00:33 +0000</pubDate>
				<category><![CDATA[Students]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[data]]></category>
		<category><![CDATA[data science]]></category>
		<category><![CDATA[machine learning]]></category>
		<guid isPermaLink="false">https://discovere.binghamton.edu/?p=7564</guid>

					<description><![CDATA[A graduate student has created an "oracle" that can make accurate predictions related to spam emails, bank fraud, workers quitting their jobs and more. ]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" class="alignleft size-full wp-image-7590" src="https://discovere.binghamton.edu/wp-content/uploads/2019/10/bani_hani_02.jpg" alt="" width="132" height="133" />Spam emails, bank fraud, diabetes, workers quitting their jobs. What do these topics have in common? The answer can be found in machine learning research at Binghamton University.</p>
<p>Dana Bani-Hani, a doctoral student studying industrial and systems engineering, has spent the past few years teaching machines how to read data sets in any industry. The system she coded, called a Recursive General Regression Neural Network Oracle (R-GRNN Oracle), takes data inputs and creates prediction outputs.</p>
<p>Classification models are not new in data science and analytics, but what Bani-Hani created goes beyond the basics. A typical system uses algorithms, called classifiers, that run through a data set of many different variables to create a prediction. Oracles are created to run multiple sets of these classifiers to see which algorithm creates the most accurate prediction.</p>
<p>For example, a classifier can look at a myriad of emails and factor in certain word usage, word count and several other variables to determine if the email is spam. An oracle looks at the different classifier outputs and determines which most accurately predicted the spam emails.</p>
<p>What sets the R-GRNN Oracle apart from other oracles is its capability to take classifier outputs and rank them based on their accuracy. Based on the ranking, classifiers are given weights and are combined to produce a prediction superior to any one classifier on its own.</p>
<p>Think of this process like an orchestra. Each instrument has its own strengths, just like different classifiers, so it is useful to include them all. The conductor, like the R-GRNN Oracle, directs the different instruments to play loudly or more softly based on how the instrument makes the final symphony sound.</p>
<p>At this point, the system would be called a General Regression Neural Network (GRNN), which has been created before at Binghamton University. The real crux of Bani-Hani’s work lies in the first letter, R, standing for Recursion.</p>
<p>The R-GRNN Oracle takes the original GRNN output, and uses that entire system as an input for another GRNN prediction. This is combined with the most successful of the original classifiers.</p>
<p>So, back to the orchestra: The original symphony is recorded, and then played back again later. This time, along with the recording, a few instruments play again to further fine-tune the important sounds of the orchestra.</p>
<p>“Because of the way [the GRNN] works, I was able to create the recursive model,” Bani-Hani says. “The concept of recursion is not widely used in machine learning, so I decided to put an oracle inside of an oracle.”</p>
<p>Mohammad Khasawneh, professor and department chair in systems science and industrial engineering, supervised Bani-Hani’s research. He says systems like the GRNN and R-GRNN are underutilized and are vital in serious life events.</p>
<p>“The traditional GRNN Oracle has received limited attention in the literature as only very few researchers have published work on the algorithm,” Khasawneh says. “But many real-life problems that apply machine learning models to automate classifying unknown observations require accurate predictions. Tasks such as diagnosing diseases entail precision to avoid serious issues that could potentially lead to problems such as lawsuits or even deaths.”</p>
<p>Bani-Hani says the R-GRNN Oracle produces more accurate predictions than any single classifier alone, as well as one GRNN on its own. The R-GRNN Oracle took in thousands of email samples, programmed to factor 57 variables, and then produced a spam prediction superior to all other classifiers tested.</p>
<p>Bani-Hani also used the R-GRNN to predict credit card application fraud, diabetes diagnosis and whether a worker will quit based on past workplace experiences. In each case, the R-GRNN came out as the most accurate predictor.</p>
<p>She plans to focus her model on specific fields, such as business or finance, as well as package both the GRNN Oracle and the R-GRNN Oracle so companies do not have to create the entire code from scratch.</p>
<p>Bani-Hani’s journey to machine learning research started nearly 6,000 miles away from Binghamton in Jordan. After completing her bachelor’s degree in architectural engineering, she heard about Binghamton University through Watson School faculty and academic leaders, and from her father’s supportive suggestions. She initially pursued a master’s degree in industrial engineering, but she soon found a new passion: data mining and machine learning.</p>
<p>&#8220;Getting a PhD has been a dream of mine for the last 15 years,&#8221; Bani-Hani says. &#8220;I mainly attribute this to having a family with advanced degrees. I am thankful to my professors here at Binghamton University for introducing me to the topics that make up my research.&#8221;</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Algorithms satisfy hunger for real-time data</title>
		<link>https://discovere.binghamton.edu/faculty-spotlights/kang-4812.html</link>
		
		<dc:creator><![CDATA[SFecht]]></dc:creator>
		<pubDate>Thu, 02 Aug 2012 07:10:40 +0000</pubDate>
				<category><![CDATA[Faculty]]></category>
		<category><![CDATA[computer science]]></category>
		<category><![CDATA[data]]></category>
		<guid isPermaLink="false">http://discovere.binghamton.edu/?p=4812</guid>

					<description><![CDATA[Computer scientist Kyoung-Don Kang's efforts to improve fast-paced data processing recently received a boost from the National Science Foundation.]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" class="alignleft size-full wp-image-4842" title="kd_kang" src="http://discovere.binghamton.edu/wp-content/uploads/2012/08/kd_kang.jpg" alt="" width="192" height="193" />The world today moves at a fast pace, and most of us don’t have time to wait around. Twitter users monitor what’s trending <em>now</em>, not last month. Drivers check the road conditions for the morning commute. Air-traffic controllers track the locations of thousands of planes simultaneously, and investors conduct high-frequency trading. Much of the data that’s collected can’t be sent to sit in a warehouse; it requires nearly instantaneous computer processing and feedback.</p>
<p>“Real-time data is very dynamic and unpredictable,” says Kyoung-Don Kang, associate professor of computer science at Binghamton University. For example, Kang explains, a traffic-monitoring system might not see much activity at midnight Sunday, but it will generate tremendous amounts of data during Monday’s morning rush hour. That data could slow down the processing system, right when it’s needed most. “It is very challenging to process this data in a timely manner,” he says.</p>
<p>When real-time computing fails, it can compromise safety or lead to financial loss.  That’s why Kang is working to make this fast-paced data processing more efficient, with help from a National Science Foundation grant of nearly $250,000.</p>
<p>“It is an important research area, especially at this point when we have lots of critical systems depending on continuous streams of real-time data from zillions of sensors deployed in the environment,” says Sang H. Son, a computer scientist at the University of Virginia.</p>
<p>Why not just design systems that are capable of processing massive amounts of data all the time? It’s not practical, Kang says, because most of the time a system will need to process only sparse amounts of data — and when it sits idle, that’s a waste of resources. And data is always increasing in volume, so even today’s top-notch system will be outpaced eventually.</p>
<p>Kang says the key to using real-time data applications is to cut your losses. If the amount of data is more than the system can handle, then some of it must be dropped, he says: “Some data is more important than others.”</p>
<p>That’s why he’ll be developing algorithms and software solutions that process the most vital data stream first. Kang will use simple yet powerful rules to prioritize some operations over others — for instance, if an input data stream is important, then the query processing output from that data stream is likely to be important as well — to build more efficient load-shedding and continuous query processing techniques. This approach can be applied to detect important events, such as unusual traffic patterns or homeland security issues, in real time.</p>
<p>More efficient processing of real-time data could one day enable other technological advances, including directing intelligent transportation or managing green buildings and smart grids. “There are many potential applications,” Kang says. “The challenge is being ready for anything.”</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
