using DataFrames
using BenchmarkToolsKnow what is copied when creating a DataFrame¶
x = DataFrame(rand(3, 5), :auto)x and y are not the same object
y = copy(x)
x === yfalsex and y are not the same object
y = DataFrame(x)
x === yfalsethe columns are also not the same
any(x[!, i] === y[!, i] for i in ncol(x))falsex and y are not the same object
y = DataFrame(x, copycols=false)
x === yfalseBut the columns are the same
all(x[!, i] === y[!, i] for i in ncol(x))truethe same when creating data frames using kwarg syntax
x = 1:3;
y = [1, 2, 3];
df = DataFrame(x=x, y=y)different object
y === df.yfalserange is converted to a vector
typeof(x), typeof(df.x)(UnitRange{Int64}, Vector{Int64})slicing rows always creates a copy
y === df[:, :y]falseyou can avoid copying by using copycols=false keyword argument in functions.
df = DataFrame(x=x, y=y, copycols=false)now it is the same
y === df.ytruenot the same object
select(df, :y)[!, 1] === yfalsethe same object
select(df, :y, copycols=false)[!, 1] === ytrueDo not modify the parent of GroupedDataFrame or view¶
x = DataFrame(id=repeat([1, 2], outer=3), x=1:6)
g = groupby(x, :id)
x[1:3, 1] = [2, 2, 2]
g ## well - it is wrong now, g is only a views = view(x, 5:6, :)delete!(x, 3:6)This is an error
s ## Will return BoundsErrorSingle column selection for a DataFrame¶
Single column selection for a DataFrame creates aliases with ! and getproperty syntax and copies with :
x = DataFrame(a=1:3)
x.b = x[!, 1] ## alias
x.c = x[:, 1] ## copy
x.d = x[!, 1][:] ## copy
x.e = copy(x[!, 1]) ## explicit copy
display(x)x[1, 1] = 100
display(x)When iterating rows of a data frame¶
use
eachrowto avoid compilation cost in wide tables,but
Tables.namedtupleiteratorfor fast execution in tall tables
The table below is tall:
df2 = DataFrame(rand(10^6, 10), :auto)@time map(sum, eachrow(df2)); 2.637511 seconds (60.12 M allocations: 1.056 GiB, 18.92% gc time, 4.89% compilation time)
@time map(sum, eachrow(df2)); 2.184906 seconds (59.99 M allocations: 1.050 GiB, 5.01% gc time)
@time map(sum, Tables.namedtupleiterator(df2)); 0.339672 seconds (1.40 M allocations: 75.099 MiB, 3.60% gc time, 98.09% compilation time)
@time map(sum, Tables.namedtupleiterator(df2)); 0.005985 seconds (21 allocations: 7.631 MiB)
as you can see - this time it is much faster to iterate a type stable container
still you might want to use the select syntax, which is optimized for such reductions:
this includes compilation time
@time select(df2, AsTable(:) => ByRow(sum) => "sum").sum 0.521406 seconds (753.65 k allocations: 45.061 MiB, 98.65% compilation time: 93% of which was recompilation)
1000000-element Vector{Float64}:
5.509193035300807
5.182497531877569
5.08255348900764
4.484686379320722
5.027928953217053
3.8628602167930435
4.39004164136988
5.22696587441208
3.6921534473356363
4.462633254974045
⋮
2.9857354718270406
5.375128642885398
6.283673085681035
5.0819468777912835
5.176628576156379
5.776846169424042
5.157416335728866
4.000946370077582
4.818008490038548Do it again
@time select(df2, AsTable(:) => ByRow(sum) => "sum").sum 0.005815 seconds (121 allocations: 7.634 MiB)
1000000-element Vector{Float64}:
5.509193035300807
5.182497531877569
5.08255348900764
4.484686379320722
5.027928953217053
3.8628602167930435
4.39004164136988
5.22696587441208
3.6921534473356363
4.462633254974045
⋮
2.9857354718270406
5.375128642885398
6.283673085681035
5.0819468777912835
5.176628576156379
5.776846169424042
5.157416335728866
4.000946370077582
4.818008490038548This notebook was generated using Literate.jl.