Hi
user
Admin Login:
Username:
Password:
Name:
Scraping Your Way to a Dataset
--client
pyohio
--show
pyohio_2019
--room cartoon1 14855 --force
Next: 12 If Statements are a Code Smell
show more...
Marks
Author(s):
Alex Zharichenko
Location
Cartoon 1
Date
jul Sat 27
Days Raw Files
Start
14:00
First Raw Start
13:32
Duration
0:45:0
Offset
0:27:21
End
14:45
Last Raw End
15:02
Chapters
00:00
0:03:13
0:33:12
Total cuts_time
41 min.
https://www.pyohio.org/2019/presentations/66
raw-playlist
raw-mp4-playlist
encoded-files-playlist
host
archive
tweet
mp4
svg
png
assets
release.pdf
Scraping_Your_Way_to_a_Dataset.json
logs
Admin:
episode
episode list
cut list
raw files day
marks day
marks day
image_files
State:
---------
borked
edit
encode
push to queue
post
richard
review 1
email
review 2
make public
tweet
to-miror
conf
done
Locked:
clear this to unlock
Locked by:
user/process that locked.
Start:
initially scheduled time from master, adjusted to match reality
Duration:
length in hh:mm:ss
Name:
Video Title (shows in video search results)
Emails:
email(s) of the presenter(s)
Released:
has someone authorised pubication
Unknown
Yes
No
Normalise:
Channelcopy:
m=mono, 01=copy left to right, 10=right to left, 00=ignore.
Thumbnail:
filename.png
Description:
markdown
It is essential to have a very large and high-quality dataset in order to perform significant analytics or to use in various machine learning tasks. For some tasks, there exists simple APIs or repositories of data to collect from. But for many other tasks like tracking prices of products, predicting stock prices, and predicting outcomes of sports games there isn't a convenient way to retrieve this information besides a webpage. Because of these circumstances, learning to scrape data from webpages and other sources allows us to create our own dataset. Additionally, scraping grants us the ability to ask better questions about data in the world. This talk is geared towards beginner-to-intermediate Python developers that want to be able to ask and answer better questions through data. This talk will provide a guide for web scraping through two examples, and it will explain how to get the scraped data into a usable form. Throughout the talk, I will highlight some tips for improving scraper performance, minimizing the risk that a web server will stop you, and different ways to store the collected data. The first of the two examples will examine a simple case of scraping data about the lottery and the second will explore a more challenging case of scraping course information from a University. Large datasets are vital for the majority of analytic and machine learning tasks. But what happens when the data you need isn't available in some convenient and easily obtainable form? This talk will go through the process of data scraping to create a dataset that can be then used for various analytical or machine learning tasks.
Comment:
production notes
2019-07-27/13_32_39.ts
Apply:
13:32:39 - 13:59:25 ( 00:26:46 )
S:
13:32:39 -
E:
14:02:38
D:
00:29:59
(
End:
1606.0)
show more...
vlc ~/Videos/veyepar/pyohio/pyohio_2019/dv/cartoon1/2019-07-27/13_32_39.ts :start-time=00.0 --audio-desync=0
Raw File
Cut List
13:32:39
seconds: 0.0
Wall: 13:32:39
Duration
00:29:59
14:02:38
seconds: 1606.0
Wall: 13:59:25
Comments:
mp4
mp4.m3u
dv.m3u
Split:
Sequence:
:
delete
2019-07-27/13_32_39.ts
Apply:
13:59:25 - 14:02:38 ( 00:03:13 )
S:
13:32:39 -
E:
14:02:38
D:
00:29:59
(
Start:
1606.0)
show more...
vlc ~/Videos/veyepar/pyohio/pyohio_2019/dv/cartoon1/2019-07-27/13_32_39.ts :start-time=01606.0 --audio-desync=0
Raw File
Cut List
13:32:39
seconds: 1606.0
Wall: 13:59:25
Duration
00:29:59
14:02:38
seconds: 0.0
Wall: 13:32:39
Comments:
mp4
mp4.m3u
dv.m3u
Split:
Sequence:
:
delete
2019-07-27/14_02_39.ts
Apply:
14:02:39 - 14:32:38 ( 00:29:59 )
S:
14:02:39 -
E:
14:32:38
D:
00:29:59
show more...
vlc ~/Videos/veyepar/pyohio/pyohio_2019/dv/cartoon1/2019-07-27/14_02_39.ts :start-time=00.0 --audio-desync=0
Raw File
Cut List
14:02:39
seconds: 0.0
Wall: 14:02:39
Duration
00:29:59
14:32:38
seconds: 0.0
Wall: 14:02:39
Comments:
mp4
mp4.m3u
dv.m3u
Split:
Sequence:
:
delete
2019-07-27/14_32_39.ts
Apply:
14:32:39 - 14:41:19 ( 00:08:40 )
S:
14:32:39 -
E:
15:02:39
D:
00:30:00
(
End:
520.0)
show more...
vlc ~/Videos/veyepar/pyohio/pyohio_2019/dv/cartoon1/2019-07-27/14_32_39.ts :start-time=00.0 --audio-desync=0
Raw File
Cut List
14:32:39
seconds: 0.0
Wall: 14:32:39
Duration
00:30:00
15:02:39
seconds: 520.0
Wall: 14:41:19
Comments:
mp4
mp4.m3u
dv.m3u
Split:
Sequence:
:
delete
2019-07-27/14_32_39.ts
Apply:
14:41:19 - 14:59:43 ( 00:18:24 )
S:
14:32:39 -
E:
15:02:39
D:
00:30:00
(
Start:
520.0) (
End:
1624.0)
show more...
vlc ~/Videos/veyepar/pyohio/pyohio_2019/dv/cartoon1/2019-07-27/14_32_39.ts :start-time=0520.0 --audio-desync=0
Raw File
Cut List
14:32:39
seconds: 520.0
Wall: 14:41:19
Duration
00:30:00
15:02:39
seconds: 1624.0
Wall: 14:59:43
Comments:
mp4
mp4.m3u
dv.m3u
Split:
Sequence:
:
delete
2019-07-27/14_32_39.ts
Apply:
14:59:43 - 15:02:39 ( 00:02:56 )
S:
14:32:39 -
E:
15:02:39
D:
00:30:00
(
Start:
1624.0)
show more...
vlc ~/Videos/veyepar/pyohio/pyohio_2019/dv/cartoon1/2019-07-27/14_32_39.ts :start-time=01624.0 --audio-desync=0
Raw File
Cut List
14:32:39
seconds: 1624.0
Wall: 14:59:43
Duration
00:30:00
15:02:39
seconds: 0.0
Wall: 14:32:39
Comments:
mp4
mp4.m3u
dv.m3u
Split:
Sequence:
:
delete
Rf filename:
root is .../show/dv/location/, example: 2013-03-13/13:13:30.dv
Sequence:
get this:
check and save to add this
2019-07-27/13_32_39.ts
2019-07-27/14_02_39.ts
2019-07-27/14_32_39.ts
Veyepar
Video Eyeball Processor and Review