{"id":3801,"date":"2020-11-18T17:53:20","date_gmt":"2020-11-18T17:53:20","guid":{"rendered":"http:\/\/ismiletechnologies.com\/?p=3801"},"modified":"2022-12-10T01:59:50","modified_gmt":"2022-12-09T20:29:50","slug":"preparing-for-serverless-big-data-open-source-software","status":"publish","type":"post","link":"https:\/\/ismiletechnologies.com\/en_us\/dataops\/preparing-for-serverless-big-data-open-source-software\/","title":{"rendered":"Preparing for serverless big data open-source software"},"content":{"rendered":"\t\t<div data-elementor-type=\"wp-post\" data-elementor-id=\"3801\" class=\"elementor elementor-3801\" data-elementor-post-type=\"post\">\n\t\t\t\t\t\t<section class=\"elementor-section elementor-top-section elementor-element elementor-element-26b4ca1 elementor-section-boxed elementor-section-height-default elementor-section-height-default\" data-id=\"26b4ca1\" data-element_type=\"section\" data-e-type=\"section\">\n\t\t\t\t\t\t<div class=\"elementor-container elementor-column-gap-default\">\n\t\t\t\t\t<div class=\"elementor-column elementor-col-100 elementor-top-column elementor-element elementor-element-7f3dde1c\" data-id=\"7f3dde1c\" data-element_type=\"column\" data-e-type=\"column\">\n\t\t\t<div class=\"elementor-widget-wrap elementor-element-populated\">\n\t\t\t\t\t\t<div class=\"elementor-element elementor-element-16006a5a elementor-widget elementor-widget-text-editor\" data-id=\"16006a5a\" data-element_type=\"widget\" data-e-type=\"widget\" data-widget_type=\"text-editor.default\">\n\t\t\t\t\t\t\t\t\tCloud technology has enabled data scientists and data analysts to deliver value without investing in an extensive infrastructure. Serverless architecture can help to reduce the associated costs of per-use billing. There are many big data open-source software such as Apache Spark, Apache Hadoop, Presto, and many others that continue to become promising in the industry-standard in enterprise data lakes and big data architecture. The below image shows the evolution of the server less future:- (Above picture)\n\nBut as they nothing comes easy and there are few challenges with the big data open-source software. With the use of server-less open-source software the separation of processing the data, more attention was given to the data proximity. The developers still have to rely on the location of the physical server to give the necessary I\/O bandwidth requirements. \u00a0All of this was done taking into consideration the available memory, I\/O characteristics, storage and compute.\n\nLet&#8217;s talk about the next phase: Serverless OSS.\u00a0But before we need to know more about the Data Proc.\n\nDATA PROC:-\u00a0Its an easy to use functional, fully managed cloud service for running managed open sources like Spark, Apache Presto, and Hadoop clusters in a simplistic manner and more cost-efficient way. This helps to cut the costs and provides advantages such as per-second pricing, idle cluster deletion, autoscaling, and much more.\n\nTo choose the ideal platform over the data application stage it&#8217;s difficult for the customers to choose from the plethora of available options and that adds to the complexity of tuning\/configuring data analytics platforms.\n\nThe complexity of tuning\/configuring data analytics platforms (processing and storage) due to the plethora of choices available to customers add to the complexity of selecting an ideal platform over the life of the data application as the usage and use case evolves. Serverless OSS will change that. \u00a0But we now have a solution and thus resolved in the steps below and there are three important aspects which can be used when delivering on QoS(Quality of Service):-\n\nCluster:- To get the desired Qos the choice of the appropriate cluster can help in the run to workload.\nInterface:- We can use the interface for workload such as (Hive, SparkSQL, Presto, Flink, and more)\nData:- It totally depends on the location, format, and data organization.\n\nIn the world of serverless data, the focus should be on the workloads and not on the infrastructure. We can do functional and automated configuration and manage the cluster which will optimize the metrics that matters to us the most, like cost and performance.\n\nIn the\u00a0server-less\u00a0world, you focus on your workloads and not on the infrastructure. We will do the automatic configuration and management of the cluster and job to\u00a0optimize\u00a0around metrics that matter to you, such as cost or performance.\t\t\t\t\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/section>\n\t\t\t\t<\/div>\n\t\t","protected":false},"excerpt":{"rendered":"<p>Cloud technology has enabled data scientists and data analysts to deliver value without investing in an extensive infrastructure. Serverless architecture can help to reduce the associated costs of per-use billing. There are many big data open-source software such as Apache Spark, Apache Hadoop, Presto, and many others that continue to become promising in the industry-standard [&hellip;]<\/p>\n","protected":false},"author":3,"featured_media":5764,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[249],"tags":[],"class_list":["post-3801","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-dataops"],"_links":{"self":[{"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/posts\/3801","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/comments?post=3801"}],"version-history":[{"count":1,"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/posts\/3801\/revisions"}],"predecessor-version":[{"id":36062,"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/posts\/3801\/revisions\/36062"}],"wp:attachment":[{"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/media?parent=3801"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/categories?post=3801"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/ismiletechnologies.com\/en_us\/wp-json\/wp\/v2\/tags?post=3801"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}